Fetching the paper…
Reading the bibliography…
Public AI benchmark results are widely broadcast by model developers as indicators of model quality within a growing and competitive market.
Construct validity in psychological tests
Lee J Cronbach and Paul E Meehl. 1955 · 1955
Earlier work this paper cites.
The Evolution of Benchmarking as a Computer Performance Evaluation Technique
Byron C. Lewis and Albert E. Crews. 1985 · 1985
Earlier work this paper cites.
Methods of coping with social desirability bias: A review
Anton J Nederhof. 1985 · 1985
Earlier work this paper cites.
Perceived Usefulness, Perceived Ease of Use, and User Acceptance of Information Technology
Fred D. Davis. 1989 · 1989
Earlier work this paper cites.
Methodology matters: Doing research in the behavioral and social sciences
Joseph E McGrath. 1995 · 1995
Earlier work this paper cites.
User Acceptance of Information Technology: Toward a Unified View
Viswanath Venkatesh, Michael G. Morris, Gordon B. Davis, and Fred D. Davis. 2003 · 2003
Earlier work this paper cites.
Developing a grounded theory approach: a comparison of Glaser and Strauss
Helen Heath and Sarah Cowley. 2004 · 2004
Earlier work this paper cites.
An empirical analysis of lead benchmarking and performance measurement: Guidance for qualitative research
Karen Anderson and Rodney McAdam. 2005 · 2005
Earlier work this paper cites.
Examining Difficulties Software Developers Encounter in The Adoption of Statistical Machine Learning. In Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 3 (Chicago, Illinois) (AAAI’08) . AAAI Press, 1563–1566
Kayur Patel, James Fogarty, James A. Landay, and Beverly Harrison. 2008 · 2008
Earlier work this paper cites.
Issues in bioinformatics benchmarking: the case study of multiple sequence alignment
Mohamed Radhouene Aniba, Olivier Poch, and Julie D. Thompson. 2010 · 2010
Earlier work this paper cites.
Fred Jelinek
Mark Liberman. 2010 · 2010
Earlier work this paper cites.
Qualitative research method: Grounded theory
Shahid N Khan. 2014 · 2014
Earlier work this paper cites.
Curiosity, creativity, and surprise as analytic tools: Grounded theory method
Michael Muller. 2014 · 2014
Earlier work this paper cites.
Conducting semi-structured interviews
William C Adams. 2015 · 2015
Earlier work this paper cites.
Basics of qualitative research . Vol. 14
Juliet Corbin and Anselm Strauss. 2015 · 2015
Cited alongside, same era.
Evaluation of Interactive Machine Learning Systems
Nadia Boukhelifa, Anastasia Bezerianos, and Evelyne Lutton. 2018 · 2018
Cited alongside, same era.
Environmental Quality Benchmarks — The Good, The Bad, and The Ugly
Peter M Chapman. 2018 · 2018
Cited alongside, same era.
Emerging trends: A tribute to Charles Wayne
Kenneth Ward Church. 2018 · 2018
Cited alongside, same era.
A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. In Proceedings of the 2020 CHI conference on human factors in computing systems . 1–12
Emma Beede, Elizabeth Baylor, Fred Hersch, Anna Iurchenko, Lauren Wilcox, Paisan Ruamviboonsuk, and Laura M Vardoulakis. 2020 · 2020
Cited alongside, same era.
Public policy and superintelligent AI: a vector field approach
AI and the Everything in the Whole Wide World Benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)
Inioluwa Deborah Raji, Emily Denton, Emily M. Bender, Alex Hanna, and Amandalynne Paullada. 2021 · 2021
Later among the works it cites.
How to Rport And Benchmark Emerging Field-Effect Transistors
Zhihui Cheng, Chin-Sheng Pang, Peiqi Wang, Son T Le, Yanqing Wu, Davood Shahrjerdi, Iuliana Radu, Max C Lemme, Lian-Mao Peng, Xiangfeng Duan, et al · 2022
Later among the works it cites.
Analyzing social settings: A guide to qualitative observation and analysis
John Lofland, David Snow, Leon Anderson, and Lyn H Lofland. 2022 · 2022
Later among the works it cites.
D2X: An eXtensible conteXtual Debugger for Modern DSLs. In Proceedings of the 21st ACM/IEEE International Symposium on Code Generation and Optimization . 162–172
Ajay Brahmakshatriya and Saman Amarasinghe. 2023 · 2023
Later among the works it cites.
A Quest through Interconnected Datasets: Research on Annotation Practices in Highly Cited Audio Machine Learning Work and Their Utilized Datasets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nick Bostrom, Allan Dafoe, and Carrick Flynn. 2020 · 2020
Cited alongside, same era.
Overcoming failures of imagination in AI infused system development and deployment
Margarita Boyarskaya, Alexandra Olteanu, and Kate Crawford. 2020 · 2020
Cited alongside, same era.
Sim2real predictivity: Does evaluation in simulation predict real-world performance?
Abhishek Kadian, Joanne Truong, Aaron Gokaslan, Alexander Clegg, Erik Wijmans, Stefan Lee, Manolis Savva, Sonia Chernova, and Dhruv Batra. 2020 · 2020
Cited alongside, same era.
Is AI Ground Truth Really True? The Dangers of Training and Evaluating AI Tools Based on Experts’ Know-What
2021 · 2021
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency . 610–623
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Cited alongside, same era.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021 · 2021
Cited alongside, same era.
Doga Tascilar. 2023 · 2023
Later among the works it cites.
Introducing the Next Generation of Claude
Anthropic. 2024 · 2024
Closest in time.
Three tensions in understanding AI–Comment on Pope Francis’ message Artificial Intelligence and the Wisdom of the Heart: Towards a Fully Human Communication
Luciano Floridi. 2024 · 2024
Closest in time.
Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation. In First Conference on Language Modeling
Declan Grabb, Max Lamparth, and Nina Vasan. 2024 · 2024
Closest in time.
ECBD: Evidence-Centered Benchmark Design for NLP
Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung, Q Vera Liao, Alexandra Olteanu, and Ziang Xiao. 2024 · 2024
Closest in time.
The AI Index 2024 Annual Report
Nestor Maslej, Loredana Fattorini, Raymond Perrault, Vanessa Parli, Anka Reuel, Erik Brynjolfsson, John Etchemendy, Katrina Ligett, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Russell Wald, and Jack Clark. 2024 · 2024
Closest in time.
GPT-4 Research and Insights
OpenAI. 2023 · 2024
Closest in time.
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
Anka Reuel*, Amelia Hardy*, Chandler Smith, Max Lamparth, and Mykel J. Kochenderfer. 2024 · 2024
Closest in time.
Benchmarks as Microscopes: A Call for Model Metrology
Michael Saxon, Ari Holtzman, Peter West, William Yang Wang, and Naomi Saphra. 2024 · 2024
Closest in time.