Fetching the paper…
Reading the bibliography…
While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged.
The Development of Intelligence in Children (The Binet-Simon Scale)
Alfred Binet and Théodore Simon · 1905
Earlier work this paper cites.
A critical examination of the concepts of face validity
Charles I Mosier · 1947
Earlier work this paper cites.
The theory and classification of criterion bias
Hubert E. Brogden and Erwin K. Taylor · 1950
Earlier work this paper cites.
Personnel selection: Test and measurement techniques
Garceau Electroencephalographs · 1952
Earlier work this paper cites.
Construct validity in psychological tests
Lee Joseph Cronbach and Paul E. Meehl · 1955
Earlier work this paper cites.
Convergent and discriminant validation by the multitrait-multimethod matrix
Donald T Campbell and Donald W Fiske · 1959
Earlier work this paper cites.
Eliza—a computer program for the study of natural language communication between man and machine
Joseph Weizenbaum · 1966
Earlier work this paper cites.
Understanding natural language
Terry Winograd · 1972
Earlier work this paper cites.
A quantitative approach to content validity
Charles Hubert Lawshe · 1975
Earlier work this paper cites.
Cognition and reality. principles and implication of cognitive psychology
Ulric Neisser · 1976
Earlier work this paper cites.
Toward an experimental ecology of human development
Urie Bronfenbrenner · 1977
Earlier work this paper cites.
Factor Analysis: Statistical Methods and Practical Issues , volume 14 of Quantitative Applications in the Social Sciences
Jae-On Kim and Charles W. Mueller · 1978
Earlier work this paper cites.
The criterion problem: 1917–1992
James T. Austin and Peter Villanova · 1992
Earlier work this paper cites.
Overview of the first trec conference
Donna Harman · 1993
Earlier work this paper cites.
Measurement invariance, factor analysis and factorial invariance
William Meredith · 1993
Earlier work this paper cites.
Chapter 9: Evaluating test validity
Lorrie A. Shepard · 1993
Earlier work this paper cites.
Constructing validity: Basic issues in objective scale development
Lee Anna Clark and David Watson · 1995
Earlier work this paper cites.
Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning
Samuel Messick · 1995
Earlier work this paper cites.
The structure of scientific revolutions , volume 962
Thomas S Kuhn · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Leon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Test validity: A matter of consequence
Samuel J. Messick · 1998
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Hidetoshi Shimodaira · 2000
Earlier work this paper cites.
Measurement validity: A shared standard for qualitative and quantitative research
Robert Adcock and David Collier · 2001
Earlier work this paper cites.
Learning a sparse representation for object detection
Shivani Agarwal and Dan Roth · 2002
Earlier work this paper cites.
Experimental and quasi-experimental designs for generalized causal inference, 2002
Charles S Reichardt · 2002
Earlier work this paper cites.
The incremental validity of psychological testing and assessment: conceptual, methodological, and statistical issues
John Hunsley and Gregory J Meyer · 2003
Earlier work this paper cites.
The concept of validity
Denny Borsboom, Gideon J Mellenbergh, and Jaap Van Heerden · 2004
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Li Fei-Fei, Rob Fergus, and Pietro Perona · 2004
Earlier work this paper cites.
Dataset issues in object recognition
Jean Ponce, Tamara L Berg, Mark Everingham, David A Forsyth, Martial Hebert, Svetlana Lazebnik, Marcin Marszalek, Cordelia Schmid, Bryan C Russell, Antonio Torralba, et al · 2006
Earlier work this paper cites.
Standard setting: A guide to establishing and evaluating performance standards on tests
Gregory J Cizek and Michael B Bunch · 2007
Earlier work this paper cites.
A suggested change in terminology and emphasis regarding validity and education
Robert W Lissitz and Karen Samuelsen · 2007
Earlier work this paper cites.
Collateral damage: How high-stakes testing corrupts america’s schools
Valerie J. Callet · 2008
Earlier work this paper cites.
Validity of the sat for predicting first-year college grade point average. research report no. 2008-5
Jennifer L Kobrin, Brian F Patterson, Emily J Shaw, Krista D Mattern, and Sandra M Barbuti · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
The weirdest people in the world?
Joseph Henrich, Steven J Heine, and Ara Norenzayan · 2010
Earlier work this paper cites.
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros · 2011
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Standards for Educational and Psychological Testing
American Educational Research Association, American Psychological Association, and National Council on Measurement in Education · 2014
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
Mark Everingham, S. M. Ali Eslami, Luc Van Gool, Christopher K. I. Williams, John M. Winn, and Andrew Zisserman · 2014
Earlier work this paper cites.
“teaching to the test” in the nclb era
Jennifer L. Jennings and Jonathan Marc Bearak · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Devi Parikh, and Dhruv Batra · 2015
Cited alongside, same era.
The ladder: A reliable leaderboard for machine learning competitions
Avrim Blum and Moritz Hardt · 2015
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Cited alongside, same era.
Experimental and Quasi-Experimental Designs for Research
Donald T. Campbell and Julian C. Stanley · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Cited alongside, same era.
On the uniform convergence of relative frequencies of events to their probabilities
Vlue: A multi-task benchmark for evaluating vision-language models
Wangchunshu Zhou, Yan Zeng, Shizhe Diao, and Xinsong Zhang · 2022
Later among the works it cites.
Impact of the covid-19 pandemic on the performance of machine learning algorithms for predicting perioperative mortality
Dimislav Ivanov Andonov, Bernhard Ulm, Martin Graessner, Armin Horst Podtschaske, Manfred Blobner, Bettina Jungwirth, and Simone Maria Kagerbauer · 2023
Later among the works it cites.
The pitfalls of untested assumptions and unwarranted/oversimplistic interpretation of cultural phenomenon: a commentary on sajjadi et al. (2023)
Mojtaba Elhami Athar · 2023
Later among the works it cites.
Fairness and machine learning: Limitations and opportunities
Solon Barocas, Moritz Hardt, and Arvind Narayanan · 2023
Later among the works it cites.
Swe-bench: Can language models resolve real-world github issues?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. N. Vapnik and A. Ya. Chervonenkis · 2015
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Cited alongside, same era.
Tackling the problem of construct proliferation
Jonathan A. Shaffer, David Scott DeGeest, and Andrew Li · 2016
Cited alongside, same era.
Findings of the 2017 conference on machine translation (wmt17)
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, and Others · 2017
Cited alongside, same era.
Structural validity of the wechsler intelligence scale for children–fifth edition: Confirmatory factor analyses with the 16 primary and secondary subtests
Gary L. Canivez, Marley W. Watkins, and Stefan C. Dombrowski · 2017
Cited alongside, same era.
An analysis of visual question answering algorithms
Kushal Kafle and Christopher Kanan · 2017
Cited alongside, same era.
Discovering causal signals in images
David Lopez-Paz, Robert Nishihara, Soumith Chintala, Bernhard Scholkopf, and Leon Bottou · 2017
Cited alongside, same era.
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Later among the works it cites.
Rethinking model evaluation as narrowing the socio-technical gap
Q Vera Liao and Ziang Xiao · 2023
Later among the works it cites.
It ain’t near ’bout fair: Re-envisioning the bias and sensitivity review process from a justice-oriented antiracist perspective
Jennifer Randall · 2023
Later among the works it cites.
It takes two to tango: Navigating conceptualizations of nlp tasks and measurements of performance
Arjun Subramonian, Xingdi Yuan, Hal Daumé III, and Su Lin Blodgett · 2023
Later among the works it cites.
From speculation to reality: Enhancing anticipatory ethics for emerging technologies (ate) in practice
Steven Umbrello, Michael J. Bernstein, Pieter E. Vermaas, Anaïs Rességuier, Gustavo Gonzalez, Andrea Porcari, Alexei Grinbaum, and Laurynas Adomaitis · 2023
Later among the works it cites.
Overwriting pretrained bias with finetuning data
Angelina Wang and Olga Russakovsky · 2023
Later among the works it cites.
Ziang Xiao, Susu Zhang, Vivian Lai, and Q Vera Liao · 2023
Later among the works it cites.
Reassessing the validity of spurious correlations benchmarks
Samuel J Bell, Diane Bouchacourt, and Levent Sagun · 2024
Later among the works it cites.
Reporting reliability, convergent and discriminant validity with structural equation modeling: A review and best-practice recommendations
Gordon W Cheung, Helena D Cooper-Thomas, Rebecca S Lau, and Linda C Wang · 2024
Later among the works it cites.
A shared standard for valid measurement of generative ai systems’ capabilities, risks, and impacts
Alexandra Chouldechova, Chad Atalla, Solon Barocas, A Feder Cooper, Emily Corvi, P Alex Dow, Jean Garcia-Gathright, Nicholas Pangakis, Stefanie Reed, Emily Sheng, et al · 2024
Later among the works it cites.
Data science at the singularity
David Donoho · 2024
Later among the works it cites.
Aryo Pradipta Gema, Haoran Liu, Yihan Diao, Jason Wei, Jiachun Liu, Donald Metzler, Shiyue Zhang, Daniel Khashabi, Jinyi Yao, Zhengbao Jiang, Chenhao Tan, and Denny Zhou · 2024
Later among the works it cites.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai
Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, et al · 2024
Later among the works it cites.
More than marketing? on the information value of ai benchmarks for practitioners
Amelia F. Hardy, Anka Reuel, K. Meimandi, Lisa Soder, Allie Griffith, Dylan M. Asmar, Sanmi Koyejo, Michael S. Bernstein, and Mykel J. Kochenderfer · 2024
Later among the works it cites.
A typology of validity: content, face, convergent, discriminant, nomological and predictive validity
Weng Marc Lim · 2024
Later among the works it cites.
Ecbd: Evidence-centered benchmark design for nlp
Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung, Q Vera Liao, Alexandra Olteanu, and Ziang Xiao · 2024
Later among the works it cites.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar · 2024
Later among the works it cites.
Ai as a sport: On the competitive epistemologies of benchmarking
Will Orr and Edward B. Kang · 2024
Later among the works it cites.
The mechanics of frictionless reproducibility
Benjamin Recht · 2024
Later among the works it cites.
Gpqa: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman · 2024
Later among the works it cites.
Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices
Anka Reuel-Lamparth, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel J Kochenderfer · 2024
Later among the works it cites.
Observational scaling laws and the predictability of language model performance
Yangjun Ruan, Chris J Maddison, and Tatsunori Hashimoto · 2024
Later among the works it cites.
Imagenot: A contrast with imagenet preserves model rankings
Olawale Salaudeen and Moritz Hardt · 2024
Later among the works it cites.
Causally inspired regularization enables domain general representations
Olawale Salaudeen and Sanmi Koyejo · 2024
Later among the works it cites.
Benchmarks as microscopes: A call for model metrology
Michael Saxon, Ari Holtzman, Peter West, William Yang Wang, and Naomi Saphra · 2024
Later among the works it cites.
Logicasker: Evaluating and improving the logical reasoning ability of large language models
Yuxuan Wan, Wenxuan Wang, Yiliu Yang, Youliang Yuan, Jen-tse Huang, Pinjia He, Wenxiang Jiao, and Michael R Lyu · 2024
Later among the works it cites.
Benchmark suites instead of leaderboards for evaluating ai fairness
Angelina Wang, Aaron Hertzmann, and Olga Russakovsky · 2024
Later among the works it cites.
Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks
Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, and Yoon Kim · 2024
Later among the works it cites.
τ \tau -bench: A benchmark for tool-agent-user interaction in real-world domains
Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan · 2024
Later among the works it cites.
Medical large language model benchmarks should prioritize construct validity
Ahmed Alaa, Thomas Hartvigsen, Niloufar Golchini, Shiladitya Dutta, Frances Dean, Inioluwa Deborah Raji, and Travis Zack · 2025
Closest in time.
Claude 3.5 sonnet
Anthropic · 2025
Closest in time.
Explicitly unbiased large language models still form biased associations
Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, and Thomas L Griffiths · 2025
Closest in time.
Ai is emerging as a research co-pilot for r&d, but can you trust it?
Brian Buntz · 2025
Closest in time.
The emerging science of machine learning benchmarks, 2025
Moritz Hardt · 2025
Closest in time.
Human-centered evaluation and auditing of language models
Yu Lu Liu, Wesley Hanwen Deng, Michelle S Lam, Motahhare Eslami, Juho Kim, Q Vera Liao, Wei Xu, Jekaterina Novikova, and Ziang Xiao · 2025
Closest in time.
NeurIPS 2024 statistics: Datasets & benchmarks track
Paper Copilot · 2025
Closest in time.
Are domain generalization benchmarks with accuracy on the line misspecified?
Olawale Salaudeen, Nicole Chiou, Shiny Weng, and Sanmi Koyejo · 2025
Closest in time.
Position: Evaluating generative ai systems is a social science measurement challenge
Hanna Wallach, Meera Desai, A Feder Cooper, Angelina Wang, Chad Atalla, Solon Barocas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P Alex Dow, et al · 2025
Closest in time.
Toward an evaluation science for generative ai systems
Laura Weidinger, Inioluwa Deborah Raji, Hanna Wallach, Margaret Mitchell, Angelina Wang, Olawale Salaudeen, Rishi Bommasani, Deep Ganguli, Sanmi Koyejo, and William Isaac · 2025
Closest in time.