Prejudice reduction: Progress and challenges
E. L. Paluck, R. Porat, C. S. Clark, and D. P. Green · 2021
Later among the works it cites.
Invalid claims about the validity of implicit association tests by prisoners of the implicit social-cognition paradigm
U. Schimmack · 2021
Later among the works it cites.
The four deadly sins of implicit attitude research
J. W. Sherman and S. A. Klein · 2021
Later among the works it cites.
Process for adapting language models to society (PALMS) with values-targeted datasets
I. Solaiman and C. Dennison · 2021
Later among the works it cites.
Multidimensional stereotypes emerge spontaneously when exploration is costly
X. Bai, T. Griffiths, and S. Fiske · 2022
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Original
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Later among the works it cites.
On the intrinsic and extrinsic fairness evaluation metrics for contextualized language representations
Y. T. Cao, Y. Pruksachatkun, K.-W. Chang, R. Gupta, V. Kumar, J. Dhamala, and A. Galstyan · 2022
Later among the works it cites.
Historical representations of social groups across 200 years of word embeddings from Google Books
T. E. Charlesworth, A. Caliskan, and M. R. Banaji · 2022
Later among the works it cites.
Holistic evaluation of language models
Original
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, et al · 2022
Later among the works it cites.
Data analysis for social science: A friendly and practical introduction
E. Llaudet and K. Imai · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Bbq: A hand-built bias benchmark for question answering
A. Parrish, A. Chen, N. Nangia, V. Padmakumar, J. Phang, J. Thompson, P. M. Htut, and S. Bowman · 2022
Later among the works it cites.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
Original
O. Shaikh, H. Zhang, W. Held, M. Bernstein, and D. Yang · 2022
Later among the works it cites.
Prompting GPT-3 to be reliable
Original
C. Si, Z. Gan, Z. Yang, S. Wang, J. Wang, J. Boyd-Graber, and L. Wang · 2022
Later among the works it cites.
Upstream mitigation is not all you need: Testing the bias transfer hypothesis in pre-trained language models
R. Steed, S. Panda, A. Kobren, and M. Wick · 2022
Later among the works it cites.
GPT-4 technical report
Original
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Later among the works it cites.
Using cognitive psychology to understand GPT-3
M. Binz and E. Schulz · 2023
Later among the works it cites.
Marked personas: Using natural language prompts to measure stereotypes in language models
M. Cheng, E. Durmus, and D. Jurafsky · 2023
Later among the works it cites.
Using large language models in psychology
D. Demszky, D. Yang, D. S. Yeager, C. J. Bryan, M. Clapper, S. Chandhok, J. C. Eichstaedt, C. Hecht, J. Jamieson, M. Johnson, et al · 2023
Later among the works it cites.
The capacity for moral self-correction in large language models
Original
D. Ganguli, A. Askell, N. Schiefer, T. I. Liao, K. Lukošiūtė, A. Chen, A. Goldie, A. Mirhoseini, andCatherine Olsson, D. Hernandez, D. Drain, D. Li, E. Tran-Johnson, E. Perez, J. Kernion, J. Kerr, J. Mueller, J. Landau, K. Ndousse, K. Nguyen, L. Lovitt, M. Sellitto, N. Elhage, N. Mercado, N. DasSarma, O. Rausch, R. Lasenby, R. Larson, S. K. Sam Ringer, S. Kadavath, S. Johnston, S. Kravec, S. E. Showk, T. Lanham, T. Telleen-Lawton, T. Henighan, T. Hume, Y. Bai, B. M. Zac Hatfield-Dodds, D. Amodei, N. Joseph, S. McCandlish, T. Brown, C. Olah, J. Clark, S. R. Bowman, and J. Kaplan · 2023
Later among the works it cites.
When do pre-training biases propagate to downstream tasks? a case study in text summarization
F. Ladhak, E. Durmus, M. Suzgun, T. Zhang, D. Jurafsky, K. McKeown, and T. Hashimoto · 2023
Later among the works it cites.
“i’m fully who i am”: Towards centering transgender and non-binary voices to measure biases in openaiopen language generation
A. Ovalle, P. Goyal, J. Dhamala, Z. Jaggers, K.-W. Chang, A. Galstyan, R. S. Zemel, and R. Gupta · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Original
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson · 2023
Later among the works it cites.
GPT is an effective tool for multilingual psychological text analysis
S. Rathje, D.-M. Mirea, I. Sucholutsky, R. Marjieh, C. Robertson, and J. J. Van Bavel · 2023
Later among the works it cites.
Evaluating and mitigating discrimination in language model decisions
Original
A. Tamkin, A. Askell, L. Lovitt, E. Durmus, N. Joseph, S. Kravec, K. Nguyen, J. Kaplan, and D. Ganguli · 2023
Later among the works it cites.
Stanford Alpaca: An instruction-following LLaMA model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Original
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Later among the works it cites.
"kelly is a warm person, joseph is a role model": Gender biases in llm-generated reference letters
Y. Wan, G. Pu, J. Sun, A. Garimella, K.-W. Chang, and N. Peng · 2023
Later among the works it cites.
Decodingtrust: A comprehensive assessment of trustworthiness in GPT models
Original
B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer, et al · 2023
Later among the works it cites.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Original
K. Zhu, J. Wang, J. Zhou, Z. Wang, H. Chen, Y. Wang, L. Yang, W. Ye, N. Z. Gong, Y. Zhang, et al · 2023
Later among the works it cites.
Dialect prejudice predicts ai decisions about people’s character, employability, and criminality
Original
V. Hofmann, P. R. Kalluri, D. Jurafsky, and S. King · 2024
Closest in time.