Fetching the paper…
Reading the bibliography…
We describe our early efforts to red team language models in order to simultaneously discover, measure, and attempt to reduce their potentially harmful outputs.
Evaluating the Underlying Gender Bias in Contextualized Word Embeddings
C. Basta, M. R. Costa-jussà, and N. Casas · 1904
Earlier work this paper cites.
Measuring Bias in Contextualized Word Representations
K. Kurita, N. Vyas, A. Pareek, A. W. Black, and Y. Tsvetkov · 1906
Earlier work this paper cites.
Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack, Aug. 2019
E. Dinan, S. Humeau, B. Chintagunta, and J. Weston · 1908
Earlier work this paper cites.
Adversarial Attacks and Defenses in Images, Graphs and Text: A Review, Oct. 2019
H. Xu, Y. Ma, H. Liu, D. Deb, H. Liu, J. Tang, and A. K. Jain · 1909
Earlier work this paper cites.
Social Bias Frames: Reasoning about Social and Power Implications of Language
M. Sap, S. Gabriel, L. Qin, D. Jurafsky, N. A. Smith, and Y. Choi · 1911
Earlier work this paper cites.
Development and validation of brief measures of positive and negative affect: the PANAS scales
D. Watson, L. A. Clark, and A. Tellegen · 1988
Earlier work this paper cites.
Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
M. Brundage, S. Avin, J. Wang, H. Belfield, G. Krueger, G. Hadfield, H. Khlaaf, J. Yang, H. Toner, R. Fong, T. Maharaj, P. W. Koh, S. Hooker, J. Leung, A. Trask, E. Bluemke, J. Lebensold, C. O’Keefe, M. Koren, T. Ryffel, J. B. Rubinovitz, T. Besiroglu, F. Carugati, J. Clark, P. Eckersley, S. de Haas, M. Johnson, B. Laurie, A. Ingerman, I. Krawczuk, A. Askell, R. Cammarota, A. Lohn, D. Krueger, C. Stix, P. Henderson, L. Graham, C. Prunkl, B. Martin, E. Seger, N. Zilberman, S. O. hÉigeartaigh, F. Kroeger, G. Sastry, R. Kagan, A. Weller, B. Tse, E. Barnes, A. Dafoe, P. Scharre, A. Herbert-Voss, M. Rasser, S. Sodhani, C. Flynn, T. K. Gilbert, L. Dyer, S. Khan, Y. Bengio, and M. Anderljung · 2004
Earlier work this paper cites.
Development and Validation of an Internationally Reliable Short-Form of the Positive and Negative Affect Schedule (PANAS)
E. R. Thompson · 2007
Earlier work this paper cites.
Can Playing the Computer Game “Tetris” Reduce the Build-Up of Flashbacks for Trauma? A Proposal from Cognitive Science
E. A. Holmes, E. L. James, T. Coode-Bate, and C. Deeprose · 2009
Earlier work this paper cites.
The Radicalization Risks of GPT-3 and Advanced Neural Language Models
K. McGuffie and A. Newhouse · 2009
Earlier work this paper cites.
New Well-being Measures: Short Scales to Assess Flourishing and Positive and Negative Feelings
E. Diener, D. Wirtz, W. Tov, C. Kim-Prieto, D.-w. Choi, S. Oishi, and R. Biswas-Diener · 2010
Earlier work this paper cites.
Extracting Training Data from Large Language Models
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, A. Oprea, and C. Raffel · 2012
Earlier work this paper cites.
Intriguing properties of neural networks, Feb. 2014
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2014
Earlier work this paper cites.
Fitting Linear Mixed-Effects Models Using lme4
D. Bates, M. Mächler, B. Bolker, and S. Walker · 2015
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Adversarial Examples for Evaluating Reading Comprehension Systems
R. Jia and P. Liang · 2017
Earlier work this paper cites.
Measuring and Mitigating Unintended Bias in Text Classification
L. Dixon, J. Li, J. Sorensen, N. Thain, and L. Vasserman · 2018
Earlier work this paper cites.
Counterfactual Fairness in Text Classification through Robustness
S. Garg, V. Perot, N. Limtiaco, A. Taly, E. H. Chi, and A. Beutel · 2019
Earlier work this paper cites.
Ghost Work
M. Gray and S. Suri · 2019
Earlier work this paper cites.
Avoiding Reasoning Shortcuts: Adversarial Evaluation, Training, and Model Development for Multi-Hop QA
Y. Jiang and M. Bansal · 2019
Earlier work this paper cites.
Testing Stylistic Interventions to Reduce Emotional Impact of Content Moderation Workers
S. Karunakaran and R. Ramakrishan · 2019
Cited alongside, same era.
But Who Protects the Moderators? The Case of Crowdsourced Image Moderation, Jan. 2020
B. Dang, M. J. Riedl, and M. Lease · 2020
Cited alongside, same era.
Fast, Accurate, and Healthier: Interactive Blurring Helps Moderators Reduce Exposure to Harmful Content
A. Das, B. Dang, and M. Lease · 2020
Cited alongside, same era.
Queens are Powerful too: Mitigating Gender Bias in Dialogue Generation
E. Dinan, A. Fan, A. Williams, J. Urbanek, D. Kiela, and J. Weston · 2020
Cited alongside, same era.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith · 2020
Cited alongside, same era.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
S. Lin, J. Hilton, and O. Evans · 2021
Later among the works it cites.
DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts
A. Liu, M. Sap, X. Lu, S. Swayamdipta, C. Bhagavatula, N. A. Smith, and Y. Choi · 2021
Later among the works it cites.
HateCheck: Functional Tests for Hate Speech Detection Models
P. Röttger, B. Vidgen, D. Nguyen, Z. Waseem, H. Margetts, and J. Pierrehumbert · 2021
Later among the works it cites.
Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets
I. Solaiman and C. Dennison · 2021
Later among the works it cites.
The Psychological Well-Being of Content Moderators: The Emotional Labor of Commercial Moderation and Avenues for Improving Support
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Social Biases in NLP Models as Barriers for Persons with Disabilities
B. Hutchinson, V. Prabhakaran, E. Denton, K. Webster, Y. Zhong, and S. Denuyl · 2020
Cited alongside, same era.
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction, Sept. 2020
L. McInnes, J. Healy, and J. Melville · 2020
Cited alongside, same era.
Adversarial NLI: A New Benchmark for Natural Language Understanding
Y. Nie, A. Williams, E. Dinan, M. Bansal, J. Weston, and D. Kiela · 2020
Cited alongside, same era.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList
M. T. Ribeiro, T. Wu, C. Guestrin, and S. Singh · 2020
Cited alongside, same era.
Large language models associate Muslims with violence
A. Abid, M. Farooqi, and J. Zou · 2021
Cited alongside, same era.
A General Language Assistant as a Laboratory for Alignment
A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. DasSarma, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, J. Kernion, K. Ndousse, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, and J. Kaplan · 2021
Cited alongside, same era.
Filling gaps in trustworthy development of AI
S. Avin, H. Belfield, M. Brundage, G. Krueger, J. Wang, A. Weller, M. Anderljung, I. Krawczuk, D. Krueger, J. Lebensold, T. Maharaj, and N. Zilberman · 2021
Cited alongside, same era.
M. Steiger, T. J. Bharucha, S. Venkatagiri, M. J. Riedl, and M. Lease · 2021
Later among the works it cites.
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
A. Tamkin, M. Brundage, J. Clark, and D. Ganguli · 2021
Later among the works it cites.
U.S. Census Bureau QuickFacts: United States, July 2021
C. US · 2021
Later among the works it cites.
Analyzing Dynamic Adversarial Training Data in the Limit, Oct. 2021
E. Wallace, A. Williams, R. Jia, and D. Kiela · 2021
Later among the works it cites.
Ethical and social risks of harm from Language Models
L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, J. Haas, L. Rimell, L. A. Hendricks, W. Isaac, S. Legassick, G. Irving, and I. Gabriel · 2021
Later among the works it cites.
Challenges in Detoxifying Language Models
J. Welbl, A. Glaese, J. Uesato, S. Dathathri, J. Mellor, L. A. Hendricks, K. Anderson, P. Kohli, B. Coppin, and P.-S. Huang · 2021
Later among the works it cites.
Bot-Adversarial Dialogue for Safe Conversational Agents
J. Xu, D. Ju, M. Li, Y.-L. Boureau, J. Weston, and E. Dinan · 2021
Later among the works it cites.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Apr. 2022
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan · 2022
Closest in time.
Predictability and Surprise in Large Generative Models
D. Ganguli, D. Hernandez, L. Lovitt, N. DasSarma, T. Henighan, A. Jones, N. Joseph, J. Kernion, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, D. Drain, N. Elhage, S. E. Showk, S. Fort, Z. Hatfield-Dodds, S. Johnston, S. Kravec, N. Nanda, K. Ndousse, C. Olsson, D. Amodei, D. Amodei, T. Brown, J. Kaplan, S. McCandlish, C. Olah, and J. Clark · 2022
Closest in time.
DALL·E 2 Preview - Risks and Limitations, 2022
P. Mishkin, L. Ahmad, M. Brundage, G. Krueger, and G. Sastry · 2022
Closest in time.
Training language models to follow instructions with human feedback, Mar. 2022
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Closest in time.
Red Teaming Language Models with Language Models
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Closest in time.
Hierarchical Text-Conditional Image Generation with CLIP Latents, Apr. 2022
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Closest in time.
LaMDA: Language Models for Dialog Applications
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, Y. Li, H. Lee, H. S. Zheng, A. Ghafouri, M. Menegali, Y. Huang, M. Krikun, D. Lepikhin, J. Qin, D. Chen, Y. Xu, Z. Chen, A. Roberts, M. Bosma, Y. Zhou, C.-C. Chang, I. Krivokon, W. Rusch, M. Pickett, K. Meier-Hellstern, M. R. Morris, T. Doshi, R. D. Santos, T. Duke, J. Soraker, B. Zevenbergen, V. Prabhakaran, M. Diaz, B. Hutchinson, K. Olson, A. Molina, E. Hoffman-John, J. Lee, L. Aroyo, R. Rajakumar, A. Butryna, M. Lamm, V. Kuzmina, J. Fenton, A. Cohen, R. Bernstein, R. Kurzweil, B. Aguera-Arcas, C. Cui, M. Croak, E. Chi, and Q. Le · 2022
Closest in time.
Adversarial Training for High-Stakes Reliability, May 2022
D. M. Ziegler, S. Nix, L. Chan, T. Bauman, P. Schmidt-Nielsen, T. Lin, A. Scherlis, N. Nabeshima, B. Weinstein-Raun, D. de Haas, B. Shlegeris, and N. Thomas · 2022
Closest in time.