Fetching the paper…
Reading the bibliography…
Quantitative Artificial Intelligence (AI) Benchmarks have emerged as fundamental tools for evaluating the performance, capability, and safety of AI models and systems.
Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher Ré · 1909
Earlier work this paper cites.
On the Measure of Intelligence, November 2019
François Chollet · 1911
Earlier work this paper cites.
Benchmarking : The Search for Industry Best Practices That Lead to Superior Performance
Robert C. Camp · 1989
Earlier work this paper cites.
"Testing - One, Two, Three … Testing!": Toward a Sociology of Testing
Trevor Pinch · 1993
Earlier work this paper cites.
‘Improving ratings’: audit in the British University system
Marilyn Strathern · 1997
Earlier work this paper cites.
Spec cpu2000: measuring cpu performance in the new millennium
J.L. Henning · 2000
Earlier work this paper cites.
Shortcut Learning in Deep Neural Networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann · 2004
Earlier work this paper cites.
From ImageNet to Image Classification: Contextualizing Progress on Benchmarks, May 2020
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas, and Aleksander Madry · 2005
Earlier work this paper cites.
Benchmarking in Optimization: Best Practice and Open Issues, December 2020
Thomas Bartz-Beielstein, Carola Doerr, Daan van den Berg, Jakob Bossek, Sowmya Chandrasekaran, Tome Eftimov, Andreas Fischbach, Pascal Kerschke, William La Cava, Manuel Lopez-Ibanez, Katherine M. Malan, Jason H. Moore, Boris Naujoks, Patryk Orzechowski, Vanessa Volz, Markus Wagner, and Thomas Weise · 2007
Earlier work this paper cites.
Bringing the People Back In: Contesting Benchmark Machine Learning Datasets, July 2020
Remi Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, Hilary Nicole, and Morgan Klaus Scheuerman · 2007
Earlier work this paper cites.
Targeting the Benchmark: On Methodology in Current Natural Language Processing Research, July 2020
David Schlangen · 2007
Earlier work this paper cites.
Utility is in the Eye of the User: A Critique of NLP Leaderboards, March 2021
Kawin Ethayarajh and Dan Jurafsky · 2009
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding, January 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2009
Earlier work this paper cites.
Issues in bioinformatics benchmarking: the case study of multiple sequence alignment
M. R. Aniba, O. Poch, and J. D. Thompson · 2010
Earlier work this paper cites.
Systematic literature studies: database searches vs. backward snowballing
Samireh Jalali and Claes Wohlin · 2012
Earlier work this paper cites.
Leakage in data mining: Formulation, detection, and avoidance
Shachar Kaufman, Saharon Rosset, Claudia Perlich, and Ori Stitelman · 2012
Earlier work this paper cites.
Data and its (dis)contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton, and Alex Hanna · 2012
Earlier work this paper cites.
Benchmarking , pages 363–368
Isabelle Bruno · 2014
Earlier work this paper cites.
Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation
Lora Aroyo and Chris Welty · 2015
Earlier work this paper cites.
Experiences from using snowballing and database searches in systematic literature studies
Deepika Badampudi, Claes Wohlin, and Kai Petersen · 2015
Earlier work this paper cites.
Turkers, Scholars, "Arafat" and "Peace": Cultural Communities and Algorithmic Gold Standards
Shilad Sen, Margaret E. Giesel, Rebecca Gold, Benjamin Hillmann, Matt Lesicko, Samuel Naden, Jesse Russell, Zixiao (Ken) Wang, and Brent Hecht · 2015
Earlier work this paper cites.
Benchmark. meaning and use, 2017
Oxford English Dictionary · 2017
Earlier work this paper cites.
Winner’s Curse? On Pace, Progress, and Empirical Rigor
D Sculley, Jasper Snoek, Ali Rahimi, and Alex Wiltschko · 2018
Earlier work this paper cites.
Another science is possible: a manifesto for slow science
Isabelle Stengers · 2018
Earlier work this paper cites.
A survey of 25 years of evaluation
Kenneth Ward Church and Joel Hestness · 2019
Earlier work this paper cites.
Model Cards for Model Reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Earlier work this paper cites.
Fairness and Abstraction in Sociotechnical Systems
Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi · 2019
Earlier work this paper cites.
How to report and benchmark emerging field-effect transistors
Z. Cheng, CS. Pang, P. Wang, and et al · 2020
Earlier work this paper cites.
Escaping the McNamara Fallacy: Toward More Impactful Recommender Systems Research
Dietmar Jannach and Christine Bauer · 2020
Earlier work this paper cites.
Put to the test: For a new sociology of testing
Noortje Marres and David Stark · 2020
Earlier work this paper cites.
Mlperf training benchmark
Peter Mattson, Christine Cheng, Gregory Diamos, Cody Coleman, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, David Brooks, Dehao Chen, Debo Dutta, Udit Gupta, Kim Hazelwood, Andy Hock, Xinyuan Huang, Daniel Kang, David Kanter, Naveen Kumar, Jeffery Liao, Deepak Narayanan, Tayo Oguntebi, Gennady Pekhimenko, Lillian Pentecost, Vijay Janapa Reddi, Taylor Robie, Tom St John, Carole-Jean Wu, Lingjie Xu, Cliff Young, and Matei Zaharia · 2020
Earlier work this paper cites.
Mlperf: An industry standard benchmark suite for machine learning performance
Peter Mattson, Vijay Janapa Reddi, Christine Cheng, Cody Coleman, Greg Diamos, David Kanter, Paulius Micikevicius, David Patterson, Guenther Schmuelling, Hanlin Tang, Gu-Yeon Wei, and Carole-Jean Wu · 2020
Earlier work this paper cites.
Evaluating Hosting Provider Security Through Abuse Data and the Creation of Metrics
Arman Noroozian · 2020
Earlier work this paper cites.
Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach · 2021
Earlier work this paper cites.
What Will it Take to Fix Benchmarking in Natural Language Understanding?
Samuel R. Bowman and George Dahl · 2021
Earlier work this paper cites.
The Benchmark Lottery, July 2021
Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals · 2021
Earlier work this paper cites.
Datasheets for Datasets, December 2021
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2021
Earlier work this paper cites.
The Constitution of Algorithms: Ground-Truthing, Programming, Formulating
Florian Jaton · 2021
Earlier work this paper cites.
Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research, December 2021
Bernard Koch, Emily Denton, Alex Hanna, and Jacob G. Foster · 2021
Earlier work this paper cites.
Question and answer test-train overlap in open-domain question answering datasets
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel · 2021
Earlier work this paper cites.
Are we learning yet? a meta review of evaluation failures across machine learning
Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt · 2021
Earlier work this paper cites.
ExplainaBoard: An Explainable Leaderboard for NLP, July 2021
Pengfei Liu, Jinlan Fu, Yang Xiao, Weizhe Yuan, Shuaicheng Chang, Junqi Dai, Yixin Liu, Zihuiwen Ye, Zi-Yi Dou, and Graham Neubig · 2021
Earlier work this paper cites.
Proxies: The Cultural Work of Standing In
Dylan Mulvin · 2021
Earlier work this paper cites.
AI and the everything in the whole wide world benchmark
Inioluwa Deborah Raji, Emily Denton, Emily M. Bender, Alex Hanna, and Amandalynne Paullada · 2021
Cited alongside, same era.
Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber · 2021
Cited alongside, same era.
“Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo · 2021
Cited alongside, same era.
Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development
Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton · 2021
Cited alongside, same era.
BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych · 2021
Cited alongside, same era.
Black-Box Access is Insufficient for Rigorous AI Audits
Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin Von Hagen, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau, Max Tegmark, David Krueger, and Dylan Hadfield-Menell · 2024
Later among the works it cites.
Algorithmic Harms and Algorithmic Wrongs
Nathalie Diberardino, Clair Baleshta, and Luke Stark · 2024
Later among the works it cites.
Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence and amending Regulations (Artificial Intelligence Act), 2024c
European Union · 2024
Later among the works it cites.
Data for Mathematical Copilots: Better Ways of Presenting Proofs for Machine Learning, December 2024
Simon Frieder, Jonas Bayer, Katherine M. Collins, Julius Berner, Jacob Loader, András Juhász, Fabian Ruehle, Sean Welleck, Gabriel Poesia, Ryan-Rhys Griffiths, Adrian Weller, Anirudh Goyal, Thomas Lukasiewicz, and Timothy Gowers · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michelle Bao, Angela Zhou, Samantha Zottola, Brian Brubach, Sarah Desmarais, Aaron Horowitz, Kristian Lum, and Suresh Venkatasubramanian · 2022
Cited alongside, same era.
Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals
Kathrin Blagec, Jakob Kraiger, Wolfgang Frühwirt, and Matthias Samwald · 2022
Cited alongside, same era.
Regulation (EU) 2022/2065 of the European Parliament and of the Council of 19 October 2022 on a Single Market For Digital Services and amending Directive 2000/31/EC (Digital Services Act), 2022
European Union · 2022
Cited alongside, same era.
Evaluation Gaps in Machine Learning Practice
Ben Hutchinson, Negar Rostamzadeh, Christina Greer, Katherine Heller, and Vinodkumar Prabhakaran · 2022
Cited alongside, same era.
Metaethical Perspectives on ’Benchmarking’ AI Ethics, April 2022
Travis LaCroix and Alexandra Sasha Luccioni · 2022
Cited alongside, same era.
Data contamination: From memorization to exploitation
Inbal Magar and Roy Schwartz · 2022
Cited alongside, same era.
What Do NLP Researchers Believe? Results of the NLP Community Metasurvey, August 2022
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, and Samuel R. Bowman · 2022
Cited alongside, same era.
Are We Done with MMLU?, June 2024
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, and Pasquale Minervini · 2024
Later among the works it cites.
Diversity in artificial intelligence conferences
Emilia Gomez, Porcaro Lorenzo, Pedro Frau Amar, and Joao Vinagre · 2024
Later among the works it cites.
Alignment faking in large language models, December 2024
Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, and Evan Hubinger · 2024
Later among the works it cites.
Constructing Capabilities: The Politics of Testing Infrastructures for Generative AI
Gabriel Grill · 2024
Later among the works it cites.
Philipp Guldimann, Alexander Spiridonov, Robin Staab, Nikola Jovanović, Mark Vero, Velko Vechev, Anna Gueorguieva, Mislav Balunović, Nikola Konstantinov, Pavol Bielik, Petar Tsankov, and Martin Vechev · 2024
Later among the works it cites.
On the Limitations of Compute Thresholds as a Governance Strategy, July 2024
Sara Hooker · 2024
Later among the works it cites.
Under the radar? examining the evaluation of foundation models
Elliot Jones, Mahi Hardalupas, and William Agrew · 2024
Later among the works it cites.
AI Agents That Matter, July 2024
Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, and Arvind Narayanan · 2024
Later among the works it cites.
Everyone Is Judging AI by These Tests. But Experts Say They’re Close to Meaningless
Jon Keegan · 2024
Later among the works it cites.
Bernard J. Koch and David Peterson · 2024
Later among the works it cites.
Questionable practices in machine learning, July 2024
Gavin Leech, Juan J. Vazquez, Niclas Kupper, Misha Yagudin, and Laurence Aitchison · 2024
Later among the works it cites.
AI competitions as infrastructures of power in medical imaging
Dieuwertje Luitse, Tobias Blanke, and Thomas Poell · 2024
Later among the works it cites.
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation, February 2024
Kristian Lum, Jacy Reese Anthis, Chirag Nagpal, and Alexander D’Amour · 2024
Later among the works it cites.
Timothy R. McIntosh, Teo Susnjak, Nalin Arachchilage, Tong Liu, Paul Watters, and Malka N. Halgamuge · 2024
Later among the works it cites.
Frontier Models are Capable of In-context Scheming, December 2024
Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn · 2024
Later among the works it cites.
State of What Art? A Call for Multi-Prompt LLM Evaluation, May 2024
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky · 2024
Later among the works it cites.
Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling, March 2024
Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba Birhane, and Inioluwa Deborah Raji · 2024
Later among the works it cites.
AI as a Sport: On the Competitive Epistemologies of Benchmarking
Will Orr and Edward B. Kang · 2024
Later among the works it cites.
Lorenzo Pacchiardi, Marko Tesic, Lucy G. Cheke, and José Hernández-Orallo · 2024
Later among the works it cites.
The Roles of English in Evaluating Multilingual Language Models, December 2024
Wessel Poelman and Miryam de Lhoneux · 2024
Later among the works it cites.
Gaps in the Safety Evaluation of Generative AI
Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Ramona Comanescu, Canfer Akbulut, Tom Stepleton, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, William Isaac, and Laura Weidinger · 2024
Later among the works it cites.
Safetywashing: Do AI safety benchmarks actually measure safety progress?
Richard Ren, Steven Basart, Adam Khoja, Alice Gatti, Long Phan, Xuwang Yin, Mantas Mazeika, Alexander Pan, Gabriel Mukobi, Ryan Hwang Kim, Stephen Fitz, and Dan Hendrycks · 2024
Later among the works it cites.
Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices
Anka Reuel, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel Kochenderfer · 2024
Later among the works it cites.
A.i. has a measurement problem
Kevin Roose · 2024
Later among the works it cites.
Paul Röttger, Fabio Pernisi, Bertie Vidgen, and Dirk Hovy · 2024
Later among the works it cites.
Lazy Data Practices Harm Fairness Research
Jan Simson, Alessandro Fabris, and Christoph Kern · 2024
Later among the works it cites.
Evaluating the World Model Implicit in a Generative Model, November 2024
Keyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan · 2024
Later among the works it cites.
AI Sandbagging: Language Models can Strategically Underperform on Evaluations, June 2024
Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F. Brown, and Francis Rhys Ward · 2024
Later among the works it cites.
Language model developers should report train-test overlap, October 2024
Andy K. Zhang, Kevin Klyman, Yifan Mai, Yoav Levine, Yian Zhang, Rishi Bommasani, and Percy Liang · 2024
Later among the works it cites.
Top llms in china and the u.s. only 5 months apart: Kai-fu lee, 2024
Lin Zhijia · 2024
Later among the works it cites.
Understanding and Benchmarking Artificial Intelligence: OpenAI’s o3 Is Not AGI, January 2025
Rolf Pfister and Hansueli Jud · 2025
Closest in time.
Framework for Artificial Intelligence Diffusion, 2025
US Department of Commerce · 2025
Closest in time.
Mapping global dynamics of benchmark creation and saturation in artificial intelligence
Simon Ott, Adriano Barbosa-Silva, Kathrin Blagec, Jan Brauner, and Matthias Samwald · 2041
Closest in time.
A noise audit of human-labeled benchmarks for machine commonsense reasoning
Mayank Kejriwal, Henrique Santos, Ke Shen, Alice M. Mulvehill, and Deborah L. McGuinness · 2045
Closest in time.
On the genealogy of machine learning datasets: A critical history of ImageNet
Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, and Hilary Nicole · 2053
Closest in time.
Agreements ‘in the wild’: Standards and alignment in machine learning benchmark dataset construction
Isak Engdahl · 2053
Closest in time.
Ground truth tracings (GTT): On the epistemic limits of machine learning
Edward B Kang · 2053
Closest in time.
Feeling fixes: Mess and emotion in algorithmic audits
Os Keyes and Jeanie Austin · 2053
Closest in time.
Politics of data reuse in machine learning systems: Theorizing reuse entanglements
Nanna Bonde Thylstrup, Kristian Bondo Hansen, Mikkel Flyverbom, and Louise Amoore · 2053
Closest in time.