Fetching the paper…
Reading the bibliography…
There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts.
Inioluwa Deborah Raji and Jingying Yang · 1912
Earlier work this paper cites.
Convergent and Discriminant Validation by the Multitrait-Multimethod Matrix
Donald T. Campbell and Donald W. Fiske · 1959
Earlier work this paper cites.
The Social Control of Technology
David Collingridge · 1982
Earlier work this paper cites.
Opinion | How law has improved auto technology
Ralph Nader · 1985
Earlier work this paper cites.
The Ontology of Complex Systems: Levels of Organization, Perspectives, and Causal Thickets
William C. Wimsatt · 1994
Earlier work this paper cites.
Pasteur’s Quadrant
Donald E. Stokes · 1997
Earlier work this paper cites.
Spirit, air, and quicksilver: The search for the "real" scale of temperature
Hasok Chang · 2001
Earlier work this paper cites.
Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes · 2001
Earlier work this paper cites.
The Gold Standard: The Challenge of Evidence-Based Medicine and Standardization in Health Care
Stefan Timmermans and Marc Berg · 2003
Earlier work this paper cites.
Beyond Accuracy: Behavioral Testing of NLP models with CheckList, May 2020
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2005
Earlier work this paper cites.
Towards Ecologically Valid Research on Language User Interfaces, July 2020
Harm de Vries, Dzmitry Bahdanau, and Christopher Manning · 2007
Earlier work this paper cites.
Proposal to develop amendments to global technical regulation No. 9 concerning pedestrian safety., 2011
United Nations · 2011
Earlier work this paper cites.
Intrinsic Bias Metrics Do Not Correlate with Application Bias, June 2021a
Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sanchez, Mugdha Pandya, and Adam Lopez · 2012
Earlier work this paper cites.
Intrinsic Bias Metrics Do Not Correlate with Application Bias, June 2021b
Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sanchez, Mugdha Pandya, and Adam Lopez · 2012
Earlier work this paper cites.
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton, and Alex Hanna · 2012
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text, October 2016
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Triangulation and the importance of establishing valid methods for food safety culture evaluation
Lone Jespersen and Carol A. Wallace · 2017
Earlier work this paper cites.
Evaluation and Reporting of Age-, Race-, and Ethnicity-Specific Data in Medical Device Clinical Studies, September 2017
U.S. Food & Drug Administration · 2017
Earlier work this paper cites.
Public policy and program evaluation
Evert Vedung · 2017
Earlier work this paper cites.
The Poison Squad: One Chemist’s Single-Minded Crusade for Food Safety at the Turn of the Twentieth Century
Deborah Blum · 2018
Earlier work this paper cites.
Diversity in Medical Device Clinical Trials: Do We Know What Works for Which Patients?
Stephanie R. Fox-Rawlings, Laura B. Gottschalk, Laurén A. Doamekpor, and Diana M. Zuckerman · 2018
Earlier work this paper cites.
Model Cards for Model Reporting, January 2019
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Earlier work this paper cites.
Machine behaviour
Iyad Rahwan, Manuel Cebrian, Nick Obradovich, Josh Bongard, Jean-François Bonnefon, Cynthia Breazeal, Jacob W. Crandall, Nicholas A. Christakis, Iain D. Couzin, Matthew O. Jackson, Nicholas R. Jennings, Ece Kamar, Isabel M. Kloumann, Hugo Larochelle, David Lazer, Richard McElreath, Alan Mislove, David C. Parkes, Alex ‘Sandy’ Pentland, Margaret E. Roberts, Azim Shariff, Joshua B. Tenenbaum, and Michael Wellman · 2019
Earlier work this paper cites.
Fairness and Abstraction in Sociotechnical Systems
Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi · 2019
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding, February 2019
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Crashworthiness
Jacobo Díaz and Miguel Costas · 2020
Earlier work this paper cites.
What Will it Take to Fix Benchmarking in Natural Language Understanding?, October 2021
Samuel R. Bowman and George E. Dahl · 2021
Earlier work this paper cites.
Datasheets for Datasets, December 2021
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2021
Earlier work this paper cites.
Dynabench: Rethinking Benchmarking in NLP, April 2021
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams · 2021
Earlier work this paper cites.
Are We Learning Yet? A Meta Review of Evaluation Failures Across Machine Learning
Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt · 2021
Earlier work this paper cites.
AI Risk Management Framework, July 2021
National Institute of Standards and Technology · 2021
Earlier work this paper cites.
The bodies underneath the rubble
Deborah Raji · 2021
Earlier work this paper cites.
AI and the Everything in the Whole Wide World Benchmark, November 2021
Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna · 2021
Cited alongside, same era.
Measuring algorithmically infused societies
Claudia Wagner, Markus Strohmaier, Alexandra Olteanu, Emre Kıcıman, Noshir Contractor, and Tina Eliassi-Rad · 2021
Cited alongside, same era.
System Safety and Artificial Intelligence, February 2022
Roel I. J. Dobbe · 2022
Cited alongside, same era.
Why Meta’s latest large language model survived only three days online, November 2022
Will Douglas Heaven · 2022
Cited alongside, same era.
Myocarditis cases reported after mRNA-based COVID-19 vaccination in the US from December 2020 to August 2021
Matthew E. Oster, David K. Shay, John R. Su, Julianne Gee, C. Buddy Creech, Karen R. Broder, Kathryn Edwards, Jonathan H. Soslow, Jeffrey M. Dendy, Elizabeth Schlaudecker, Sean M. Lang, Elizabeth D. Barnett, Frederick L. Ruberg, Michael J. Smith, M. Jay Campbell, Renato D. Lopes, Laurence S. Sperling, Jane A. Baumblatt, Deborah L. Thompson, Paige L. Marquez, Penelope Strid, Jared Woo, River Pugsley, Sarah Reagan-Steiner, Frank DeStefano, and Tom T. Shimabukuro · 2022
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, December 2024
Gemini, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, Andrea Tacchetti, Colin Gaffney, Samira Daruki, Olcan Sercinoglu, Zach Gleicher, Juliette Love, Paul Voigtlaender, Rohan Jain, Gabriela Surita, Kareem Mohamed, Rory Blevins, Junwhan Ahn, Tao Zhu, Kornraphop Kawintiranon, Orhan Firat, Yiming Gu, Yujing Zhang, Matthew Rahtz, Manaal Faruqui, Natalie Clay, Justin Gilmer, J. D. Co-Reyes, Ivo Penchev, Rui Zhu, Nobuyuki Morioka, Kevin Hui, Krishna Haridasan, Victor Campos, Mahdis Mahdieh, Mandy Guo, and Samer Hassan et al · 2024
Later among the works it cites.
The Llama 3 Herd of Models, November 2024
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, Danny Wyatt, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, and Ehab AlBadawy et al · 2024
Later among the works it cites.
Open-Endedness is Essential for Artificial Superhuman Intelligence, June 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Challenges in evaluating AI systems, October 2023
Anthropic · 2023
Cited alongside, same era.
Microsoft and Epic expand AI collaboration to accelerate generative AI’s impact in healthcare, addressing the industry’s most pressing needs, August 2023
Eric Boyd · 2023
Cited alongside, same era.
What the data says about Americans’ views of artificial intelligence, November 2023
Michelle Faverio and Alec Tyson · 2023
Cited alongside, same era.
AI Red-Teaming Is Not a One-Stop Solution to AI Harms: Recommendations for Using Red-Teaming for AI Accountability, 2023
Sorelle Friedler, Ranjit Singh, Borhane Blili-Hamelin, Jacob Metcalfe, and Brian J. Chen · 2023
Cited alongside, same era.
ChatGPT sets record for fastest-growing user base - analyst note, February 2023
Krystal Hu · 2023
Cited alongside, same era.
Towards a Science of Human-AI Decision Making: An Overview of Design Space in Empirical Human-Subject Studies
Vivian Lai, Chacha Chen, Alison Smith-Renner, Q. Vera Liao, and Chenhao Tan · 2023
Cited alongside, same era.
AI safety on whose terms?
Seth Lazar and Alondra Nelson · 2023
Cited alongside, same era.
Edward Hughes, Michael Dennis, Jack Parker-Holder, Feryal Behbahani, Aditi Mavalankar, Yuge Shi, Tom Schaul, and Tim Rocktaschel · 2024
Later among the works it cites.
Anna Kawakami, Daricia Wilkinson, and Alexandra Chouldechova · 2024
Later among the works it cites.
What’s documented in AI? Systematic Analysis of 32K AI Model Cards, February 2024
Weixin Liang, Nazneen Rajani, Xinyu Yang, Ezinwanne Ozoani, Eric Wu, Yiqun Chen, Daniel Scott Smith, and James Zou · 2024
Later among the works it cites.
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild, October 2024
Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Faeze Brahman, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, and Yejin Choi · 2024
Later among the works it cites.
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation, February 2024
Kristian Lum, Jacy Reese Anthis, Chirag Nagpal, and Alexander D’Amour · 2024
Later among the works it cites.
AI on trial: legal models hallucinate in 1 out of 6 (or more) benchmarking queries, May 2024
Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning, and Daniel E. Ho · 2024
Later among the works it cites.
The Code That Binds Us: Navigating the Appropriateness of Human-AI Assistant Relationships
Arianna Manzini, Geoff Keeling, Lize Alberts, Shannon Vallor, Meredith Ringel Morris, and Iason Gabriel · 2024
Later among the works it cites.
Embers of autoregression show how large language models are shaped by the problem they are trained to solve
R. Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D. Hardy, and Thomas L. Griffiths · 2024
Later among the works it cites.
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar · 2024
Later among the works it cites.
The way we measure progress in AI is terrible
Scott J. Mulligan · 2024
Later among the works it cites.
Findings & recommendations: AI safety
The National Artificial Intelligence Advisory Committee NAIAC · 2024
Later among the works it cites.
How Beginning Programmers and Code LLMs (Mis)read Each Other, 2024
Sydney Nguyen, Hannah McLean Babe, Yangtian Zi, Arjun Guha, Carolyn Jane Anderson, and Molly Q Feldman · 2024
Later among the works it cites.
Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling, March 2024
Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba Birhane, and Inioluwa Deborah Raji · 2024
Later among the works it cites.
GPT-4o System Card, 2024
OpenAI · 2024
Later among the works it cites.
Gaps in the Safety Evaluation of Generative AI
Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Ramona Comanescu, Canfer Akbulut, Tom Stepleton, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, William Isaac, and Laura Weidinger · 2024
Later among the works it cites.
Generative AI in Search: Let Google do the searching for you, May 2024
Liz Reid · 2024
Later among the works it cites.
A.I. Has a Measurement Problem
Kevin Roose · 2024
Later among the works it cites.
Can A.I. Be Blamed for a Teen’s Suicide?
Kevin Roose · 2024
Later among the works it cites.
Benchmarks as Microscopes: A Call for Model Metrology, July 2024
Michael Saxon, Ari Holtzman, Peter West, William Yang Wang, and Naomi Saphra · 2024
Later among the works it cites.
Evaluating the Social Impact of Generative AI Systems in Systems and Society, June 2024
Irene Solaiman, Zeerak Talat, William Agnew, Lama Ahmad, Dylan Baker, Su Lin Blodgett, Canyu Chen, Hal Daumé III, Jesse Dodge, Isabella Duan, Ellie Evans, Felix Friedrich, Avijit Ghosh, Usman Gohar, Sara Hooker, Yacine Jernite, Ria Kalluri, Alberto Lusoli, Alina Leidinger, Michelle Lin, Xiuzhu Lin, Sasha Luccioni, Jennifer Mickel, Margaret Mitchell, Jessica Newman, Anaelia Ovalle, Marie-Therese Png, Shubham Singh, Andrew Strait, Lukas Struppek, and Arjun Subramonian · 2024
Later among the works it cites.
Evaluating Generative AI Systems is a Social Science Measurement Challenge, November 2024
Hanna Wallach, Meera Desai, Nicholas Pangakis, A. Feder Cooper, Angelina Wang, Solon Barocas, Alexandra Chouldechova, Chad Atalla, Su Lin Blodgett, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Alexandra Olteanu, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vaughan, Matthew Vogel, Hannah Washington, and Abigail Z. Jacobs · 2024
Later among the works it cites.
Benchmark suites instead of leaderboards for evaluating ai fairness
Angelina Wang, Aaron Hertzmann, and Olga Russakovsky · 2024
Later among the works it cites.
The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark, September 2024
Kyle Wiggers · 2024
Later among the works it cites.
A Careful Examination of Large Language Model Performance on Grade School Arithmetic, November 2024
Hugh Zhang, Jeff Da, Dean Lee, Vaughn Robinson, Catherine Wu, Will Song, Tiffany Zhao, Pranav Raja, Charlotte Zhuang, Dylan Slack, Qin Lyu, Sean Hendryx, Russell Kaplan, Michele Lunati, and Summer Yue · 2024
Later among the works it cites.
HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions, October 2024
Xuhui Zhou, Hyunwoo Kim, Faeze Brahman, Liwei Jiang, Hao Zhu, Ximing Lu, Frank Xu, Bill Yuchen Lin, Yejin Choi, Niloofar Mireshghallah, Ronan Le Bras, and Maarten Sap · 2024
Later among the works it cites.
Representation of BBC News content in AI Assistants
Pete Archer and Oli Elliott · 2025
Closest in time.
Not All Clinical AI Monitoring Systems Are Created Equal: Review and Recommendations
Jean Feng, Fan Xia, Karandeep Singh, and Romain Pirracchio · 2025
Closest in time.
Nari Johnson, Elise Silva, Harrison Leon, Motahhare Eslami, Beth Schwanke, Ravit Dotan, and Hoda Heidari · 2025
Closest in time.
Removing Barriers to American Leadership in Artificial Intelligence, January 2025
White House · 2025
Closest in time.