Fetching the paper…
Reading the bibliography…
This paper introduces RiskCards, a framework for structured assessment and documentation of risks associated with an application of language models.
Risk Analysis Systems for Veterinary Biologicals: A Regulator’s Tool Box
C. G. Osborne, M. D. McElvaine, A. S. Ahl, and J. W. Glosser. 1995 · 1995
Earlier work this paper cites.
Emancipatory research: Realistic goal or impossible dream
Michael Oliver. 1997 · 1997
Earlier work this paper cites.
The methodology of participatory design
Clay Spinuzzi. 2005 · 2005
Earlier work this paper cites.
The Character of Harms: Operational Challenges in Control
Malcolm K. Sparrow. 2008 · 2008
Earlier work this paper cites.
Case C-343/09 Afton Chemical Limited v Secretary of State for Transport
European Court of Justice. 2010 · 2010
Earlier work this paper cites.
Understanding Regulation: Theory, Strategy, and Practice (2nd ed ed.)
Robert Baldwin, Martin Cave, and Martin Lodge. 2012 · 2012
Earlier work this paper cites.
Evidencing the harms of hate speech
Katharine Gelber and Luke McNamara. 2016 · 2016
Earlier work this paper cites.
Hate speech in Rwanda: The road to genocide
William A Schabas. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
What is the harm of hate speech?
Eric Barendt. 2019 · 2019
Earlier work this paper cites.
Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency . 220–229
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019 · 2019
Earlier work this paper cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2019 · 2019
Earlier work this paper cites.
The Risk of Racial Bias in Hate Speech Detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Florence, Italy, 1668–1678
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019 · 2019
Earlier work this paper cites.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al · 2019
Earlier work this paper cites.
Violating Rights: Enforcing the World’s Blasphemy Laws
USCIRF. 2019 · 2019
Earlier work this paper cites.
Trick me if you can: Human-in-the-loop generation of adversarial examples for question answering
Eric Wallace, Pedro Rodriguez, Shi Feng, Ikuya Yamada, and Jordan Boyd-Graber. 2019 · 2019
Earlier work this paper cites.
A unified taxonomy of harmful content. In Proceedings of the fourth Workshop on Online Abuse and Harms . Association for Computational Linguistics, 125–137
Michele Banko, Brendon MacKeen, and Laurie Ray. 2020 · 2020
Earlier work this paper cites.
Beat the AI: Investigating adversarial human annotation for reading comprehension
Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp. 2020 · 2020
Cited alongside, same era.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2020 . Association for Computational Linguistics, Online, 3356–3369
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
Arguments for an ‘emancipatory’research paradigm
Beth Humphries, Donna M Mertens, and Carole Truman. 2020 · 2020
Cited alongside, same era.
The State and Fate of Linguistic Diversity and Inclusion in the NLP World. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Online, 6282–6293
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Cited alongside, same era.
Towards generalisable hate speech detection: a review on obstacles and solutions
Wenjie Yin and Arkaitz Zubiaga. 2021 · 2021
Later among the works it cites.
Machine Generated Text: A Comprehensive Survey of Threat Models and Detection Methods
Evan Crothers, Nathalie Japkowicz, and Herna Viktor. 2022 · 2022
Later among the works it cites.
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Later among the works it cites.
Handling and Presenting Harmful Text in NLP Research. In Findings of the Association for Computational Linguistics: EMNLP 2022
Hannah Rose Kirk, Abeba Birhane, Bertie Vidgen, and Leon Derczynski. 2022a · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Franziska B Keller, David Schoch, Sebastian Stier, and JungHwan Yang. 2020 · 2020
Cited alongside, same era.
Towards a comprehensive taxonomy and large-scale annotated corpus for online slur usage. In Proceedings of the Fourth Workshop on Online Abuse and Harms . 138–149
Jana Kurrek, Haji Mohammad Saleem, and Derek Ruths. 2020 · 2020
Cited alongside, same era.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2020 · 2020
Cited alongside, same era.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman. 2020 · 2020
Cited alongside, same era.
Red Team Development and Operations–A practical Guide
Joe Vest and James Tubberville. 2020 · 2020
Cited alongside, same era.
Directions in abusive language training data, a systematic review: Garbage in, garbage out
Bertie Vidgen and Leon Derczynski. 2020 · 2020
Cited alongside, same era.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2020 · 2020
Cited alongside, same era.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . ACM, 610–623
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Hannah Rose Kirk, Bertram Vidgen, Paul Röttger, Tristan Thrush, and Scott A. Hale. 2022b · 2022
Later among the works it cites.
Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable Survey
Sachin Kumar, Vidhisha Balachandran, Lucille Njoo, Antonios Anastasopoulos, and Yulia Tsvetkov. 2022 · 2022
Later among the works it cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Later among the works it cites.
DALL·E 2 Preview - Risks and Limitations
P Mishkin, L Ahmad, M Brundage, G Krueger, and G Sastry. 2022 · 2022
Later among the works it cites.
Ethics Sheets for AI Tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Dublin, Ireland, 8368–8379
Saif Mohammad. 2022 · 2022
Later among the works it cites.
Jailbreaking ChatGPT on Release Day
Zvi Mowshowitz. 2022 · 2022
Later among the works it cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Later among the works it cites.
Sociotechnical Harms: Scoping a Taxonomy for Harm Reduction
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla, Jess Gallegos, Andrew Smart, Emilio Garcia, et al · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Later among the works it cites.
Building Human Values into Recommender Systems: An Interdisciplinary Synthesis
Jonathan Stray, Alon Halevy, Parisa Assar, Dylan Hadfield-Menell, Craig Boutilier, Amar Ashar, Lex Beattie, Michael Ekstrand, Claire Leibowicz, Connie Moon Sehat, et al · 2022
Later among the works it cites.
Reverse Prompt Engineering for Fun and (no) Profit
Sean Wang. 2022 · 2022
Later among the works it cites.
Taxonomy of risks posed by language models. In 2022 ACM Conference on Fairness, Accountability, and Transparency . 214–229
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al · 2022
Later among the works it cites.
Regulating the Risks of AI
Margot E Kaminski. 2023 · 2023
Closest in time.
Auditing large language models: a three-layered approach
Jakob Mökander, Jonas Schuett, Hannah Rose Kirk, and Luciano Floridi. 2023 · 2023
Closest in time.