Fetching the paper…
Reading the bibliography…
Benchmark datasets play a central role in the organization of machine learning research.
The measurement of diversity in different types of biological collections
Evelyn C Pielou · 1966
Earlier work this paper cites.
The Sociology of Science: Theoretical and Empirical Investigations
R.K. Merton and N.W. Storer · 1973
Earlier work this paper cites.
Why the social sciences won’t become high-consensus, rapid-discovery science
Randall Collins · 1994
Earlier work this paper cites.
A theory of benchmarking with applications to software reverse engineering
Susan Elliott Sim · 2003
Earlier work this paper cites.
Using benchmarking to advance research: a challenge to software engineering
S.E. Sim, S. Easterbrook, and R.C. Holt · 2003
Earlier work this paper cites.
The small-sample bias of the gini coefficient: results and implications for empirical research
George Deltas · 2003
Earlier work this paper cites.
A better lemon squeezer? maximum-likelihood regression with beta-distributed dependent variables
Michael Smithson and Jay Verkuilen · 2006
Earlier work this paper cites.
The Generalized Beta Distribution as a Model for the Distribution of Income: Estimation of Related Measures of Inequality
James B. McDonald and Michael Ransom · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The megaface benchmark: 1 million faces for recognition at scale
Ira Kemelmacher-Shlizerman, Steven M Seitz, Daniel Miller, and Evan Brossard · 2016
Earlier work this paper cites.
Ms-celeb-1m: A dataset and benchmark for large-scale face recognition
Yandong Guo, Lei Zhang, Yuxiao Hu, Xiaodong He, and Jianfeng Gao · 2016
Earlier work this paper cites.
The perpetual line-up: Unregulated police face recognition in America
Clare Garvie, Alvaro Bedoya, and Jonathan Frankle · 2016
Earlier work this paper cites.
No classification without representation: Assessing geodiversity issues in open data sets for the developing world
S. Shankar, Yoni Halpern, Eric Breck, J. Atwood, Jimbo Wilson, and D. Sculley · 2017
Earlier work this paper cites.
Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman · 2018
Cited alongside, same era.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2018
Cited alongside, same era.
Emerging trends: A tribute to Charles Wayne
Kenneth Ward Church · 2018
Cited alongside, same era.
AI and compute
Dario Amodei, Danny Hernandez, Girish Sastry, Jack Clark, Greg Brockman, and Ilya Sutskever · 2018
Cited alongside, same era.
Google’s AI guru wants computers to think more like brains
Tom Simonite · 2018
Cited alongside, same era.
Winner’s curse? On pace, progress, and empirical rigor
D. Sculley, Jasper Snoek, Alexander B. Wiltschko, and A. Rahimi · 2018
Data and its (dis)contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton, and Alex Hanna · 2020
Later among the works it cites.
Garbage in, garbage out? do machine learning application papers in social computing report where human-labeled training data comes from?
R. Stuart Geiger, Kevin Yu, Yanlai Yang, Mindy Dai, Jie Qiu, Rebekah Tang, and Jenny Huang · 2020
Later among the works it cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Later among the works it cites.
Underspecification presents challenges for credibility in modern machine learning
Alexander D’Amour, Katherine A. Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yi-An Ma, Cory Y. McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai, and D. Sculley · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Excavating AI: The politics of images in machine learning training sets
Kate Crawford and Trevor Paglen · 2019
Cited alongside, same era.
A survey of 25 years of evaluation
Kenneth Ward Church and Joel Hestness · 2019
Cited alongside, same era.
Green AI
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni · 2019
Cited alongside, same era.
Show your work: Improved reporting of experimental results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith · 2019
Cited alongside, same era.
NLP’s Clever Hans moment has arrived
Benjamin Heinzerling · 2019
Cited alongside, same era.
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila · 2019
Cited alongside, same era.
AI and the everything in the whole wide world benchmark
Deborah I Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna · 2020
Later among the works it cites.
Utility is in the eye of the user: A critique of nlp leaderboards
Kawin Ethayarajh and Dan Jurafsky · 2020
Later among the works it cites.
Are we done with ImageNet?
L. Beyer, Olivier J. H’enaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord · 2020
Later among the works it cites.
From ImageNet to image classification: Contextualizing progress on benchmarks
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas, and Aleksander Madry · 2020
Later among the works it cites.
Microsoft academic graph: When experts are not enough
Kuansan Wang, Zhihong Shen, Chiyuan Huang, Chieh-Han Wu, Yuxiao Dong, and Anshul Kanakia · 2020
Later among the works it cites.
Large image datasets: A pyrrhic win for computer vision?
Abeba Birhane and Vinay Uday Prabhu · 2021
Closest in time.
Do datasets have politics? Disciplinary values in computer vision dataset development
Morgan Klaus Scheuerman, Emily Denton, and Alex Hanna · 2021
Closest in time.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Closest in time.
"Everyone wants to do the model work, not the data work": Data cascades in high-stakes AI"
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Kumar Paritosh, and Lora Mois Aroyo · 2021
Closest in time.
Exposing AI
Adam Harvey and Jules LaPlace · 2021
Closest in time.
Science on a Mission: How Military Funding Shaped What We Do and Don’t Know about the Ocean
Naomi Oreskes · 2021
Closest in time.