Fetching the paper…
Reading the bibliography…
In recent years, ML researchers have wrestled with defining and improving machine learning (ML) benchmarks and datasets.
"General Intelligence," Objectively Determined and Measured
C. Spearman. 1904 · 1904
Earlier work this paper cites.
A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, Andre Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, Lucas Beyer, Olivier Bachem, Michael Tschannen, Marcin Michalski, Olivier Bousquet, Sylvain Gelly, and Neil Houlsby. 2019 · 1910
Earlier work this paper cites.
Der Sinn der “Wertfreiheit“ der soziologischen und ökonomischen Wissenschaften
Max Weber. 1917 · 1917
Earlier work this paper cites.
Objectivity, Value Judgment, and Theory Choice
Thomas S. Kuhn. 1977 · 1977
Earlier work this paper cites.
Autonomous technology: technics-out-of-control as a theme in political thought
Langdon Winner. 1978 · 1978
Earlier work this paper cites.
The mismeasure of man (1st ed.)
Stephen Jay Gould. 1981 · 1981
Earlier work this paper cites.
Straight talk about mental tests
Arthur Robert Jensen. 1981 · 1981
Earlier work this paper cites.
Human cognitive abilities: a survey of factor-analytic studies
John B. Carroll. 1993 · 1993
Earlier work this paper cites.
Path dependence, lock-in, and history
Stan J Liebowitz and Stephen E Margolis. 1995 · 1995
Earlier work this paper cites.
Cognitive and Non-Cognitive Values in Science: Rethinking the Dichotomy
Helen E. Longino. 1996 · 1996
Earlier work this paper cites.
Why g matters: The complexity of everyday life
Linda S. Gottfredson. 1997 · 1997
Earlier work this paper cites.
The Black–White test score gap
Christopher Jencks and Meredith Phillips (Eds.). 1998 · 1998
Earlier work this paper cites.
The g factor: the science of mental ability
Arthur Robert Jensen. 1998 · 1998
Earlier work this paper cites.
Sorting things out: Classification and its consequences
Geoffrey C Bowker and Susan Leigh Star. 2000 · 2000
Earlier work this paper cites.
Inductive Risk and Values in Science
Heather Douglas. 2000 · 2000
Earlier work this paper cites.
Situated Knowledge and the Interplay of Value Judgments and Evidence in Scientific Inquiry
Elizabeth Anderson. 2002 · 2002
Earlier work this paper cites.
Premorbid cognitive testing predicts the onset of dementia and Alzheimer’s disease better than and independently of APOE genotype
J Cervilla. 2004 · 2003
Earlier work this paper cites.
Teachers’ perceptions and expectations and the Black-White test score gap
Ronald F Ferguson. 2003 · 2003
Earlier work this paper cites.
Intelligence Predicts Health and Longevity, but Why?
Linda S. Gottfredson and Ian J. Deary. 2004 · 2004
Earlier work this paper cites.
Intergenerational social mobility and mid-life status attainment: Influences of childhood intelligence, childhood social factors, and education
Ian J. Deary, Michelle D. Taylor, Carole L. Hart, Valerie Wilson, George Davey Smith, David Blane, and John M. Starr. 2005 · 2005
Earlier work this paper cites.
Intelligence and educational achievement
Ian J. Deary, Steve Strand, Pauline Smith, and Cres Fernandes. 2007 · 2006
Earlier work this paper cites.
Intelligence and socioeconomic success: A meta-analytic review of longitudinal research
Tarmo Strenze. 2007 · 2006
Earlier work this paper cites.
Bringing the People Back In: Contesting Benchmark Machine Learning Datasets
Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, Hilary Nicole, and Morgan Klaus Scheuerman. 2020 · 2007
Earlier work this paper cites.
Wechsler Adult Intelligence Scale: WAIS-IV ; technical and interpretive manual (4th ed.)
David Wechsler. 2008 · 2008
Earlier work this paper cites.
Path Dependence in the Production of Scientific Knowledge
Mark S. Peacock. 2009 · 2009
Earlier work this paper cites.
Epistemic Values and the Argument from Inductive Risk
Daniel Steel. 2010 · 2010
Earlier work this paper cites.
Overcoming Failures of Imagination in AI Infused System Development and Deployment
Margarita Boyarskaya, Alexandra Olteanu, and Kate Crawford. 2020 · 2011
Earlier work this paper cites.
The Meaning of “Ethical Neutrality” in Sociology and Economics
Max Weber. 2011 · 2011
Earlier work this paper cites.
125 Years of Intelligence in The American Journal of Psychology
Ian J. Deary. 2012a · 2012
Earlier work this paper cites.
Intelligence
Ian J. Deary. 2012b · 2012
Earlier work this paper cites.
Is chess the drosophila of artificial intelligence? A social history of an algorithm
Nathan Ensmenger. 2012 · 2012
Earlier work this paper cites.
Values in Science
Ernan McMullin. 1982 · 2012
Earlier work this paper cites.
Introduction: Thick and Thin Concepts
Simon Kirchin. 2013 · 2013
Earlier work this paper cites.
The Moral Terrain of Science
Heather Douglas. 2014 · 2014
Earlier work this paper cites.
Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation
Lora Aroyo and Chris Welty. 2015 · 2015
Earlier work this paper cites.
Intelligence: all that matters
Stuart Ritchie. 2015 · 2015
Earlier work this paper cites.
The association between intelligence and lifespan is mostly genetic
Rosalind Arden, Michelle Luciano, Ian J Deary, Chandra A Reynolds, Nancy L Pedersen, Brenda L Plassman, Matt McGue, Kaare Christensen, and Peter M Visscher. 2016 · 2016
Earlier work this paper cites.
Where are human subjects in Big Data research? The emerging ethics divide
Jacob Metcalf and Kate Crawford. 2016 · 2016
Earlier work this paper cites.
Five Reasons to Put the g
Russell T. Warne. 2016 · 2016
Cited alongside, same era.
Intelligence, Disability, and Race: Intersections and Critical Questions
Licia Carlson. 2017 · 2017
Cited alongside, same era.
Thick evaluation (1st ed.)
Simon Kirchin. 2017 · 2017
Cited alongside, same era.
Eugenics: a very short introduction
Philippa Levine. 2017 · 2017
Cited alongside, same era.
Innate: how the wiring of our brains shapes who we are
Kevin J. Mitchell. 2018 · 2018
Cited alongside, same era.
Intelligence Is an Ableist Concept by Amy Sequenzia on Ollibean
Amy Sequenzia. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Institutionalizing ethics in AI through broader impact requirements
Carina E. A. Prunkl, Carolyn Ashurst, Markus Anderljung, Helena Webb, Jan Leike, and Allan Dafoe. 2021 · 2021
Later among the works it cites.
AI and the everything in the whole wide world benchmark
Inioluwa Deborah Raji, Emily M Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. 2021 · 2021
Later among the works it cites.
Do datasets have politics? Disciplinary values in computer vision dataset development
Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton. 2021 · 2021
Later among the works it cites.
The ethics of emotion in artificial intelligence systems. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . 782–793
Luke Stark and Jesse Hoey. 2021 · 2021
Later among the works it cites.
Consequences, Schmonsequences! Considering the Future as Part of Publication and Peer Review in Computing Research. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems . ACM, Yokohama Japan, 1–4
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Cited alongside, same era.
Leveling the playing field: Fairness in AI versus human game benchmarks. In Proceedings of the 14th International Conference on the Foundations of Digital Games . 1–8
Rodrigo Canaan, Christoph Salge, Julian Togelius, and Andy Nealen. 2019 · 2019
Cited alongside, same era.
Show Your Work: Improved Reporting of Experimental Results. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . Association for Computational Linguistics, Hong Kong, China, 2185–2194
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Key challenges for delivering clinical impact with artificial intelligence
Christopher J Kelly, Alan Karthikesalingam, Mustafa Suleyman, Greg Corrado, and Dominic King. 2019 · 2019
Cited alongside, same era.
How Well Do Machines Perform on IQ tests: a Comparison Study on a Large-Scale Dataset. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence . International Joint Conferences on Artificial Intelligence Organization, Macao, China, 6110–6116
Yusen Liu, Fangyuan He, Haodi Zhang, Guozheng Rao, Zhiyong Feng, and Yi Zhou. 2019 · 2019
Cited alongside, same era.
Rebooting AI: Building artificial intelligence we can trust
Gary Marcus and Ernest Davis. 2019 · 2019
Cited alongside, same era.
Miriam Sturdee, Joseph Lindley, Conor Linehan, Chris Elsden, Neha Kumar, Tawanna Dillahunt, Regan Mandryk, and John Vines. 2021 · 2021
Later among the works it cites.
Thick Ethical Concepts
Pekka Väyrynen. 2021 · 2021
Later among the works it cites.
Between-Group Mean Differences in Intelligence in the United States Are >0% Genetically Caused: Five Converging Lines of Evidence
Russell T. Warne. 2021 · 2021
Later among the works it cites.
Democratising Measurement: or Why Thick Concepts Call for Coproduction
Anna Alexandrova and Mark Fabian. 2022 · 2022
Closest in time.
Provisional Draft of the NeurIPS Code of Ethics
Samy Bengio, Alina Beygelzimer, Kate Crawford, Jeanne Fromer, Iason Gabriel, Amanda Levendowski, Deborah Raji, and Marc’Aurelio Ranzato. 2022 · 2022
Closest in time.
The Values Encoded in Machine Learning Research. In 2022 ACM Conference on Fairness, Accountability, and Transparency . ACM, Seoul, 173–184
Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. 2022 · 2022
Closest in time.
Evaluation for Change
Rishi Bommasani. 2022 · 2022
Closest in time.
Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?
Rishi Bommasani, Kathleen A. Creel, Ananya Kumar, Dan Jurafsky, and Percy Liang. 2022 · 2022
Closest in time.
Emerging trends: SOTA-chasing
Kenneth Ward Church and Valia Kordoni. 2022 · 2022
Closest in time.
A Validity Perspective on Evaluating the Justified Use of Data-driven Decision-making Algorithms
Amanda Coston, Anna Kawakami, Haiyi Zhu, Ken Holstein, and Hoda Heidari. 2022 · 2022
Closest in time.
The Algorithmic Leviathan: Arbitrariness, Fairness, and Opportunity in Algorithmic Decision-Making Systems
Kathleen Creel and Deborah Hellman. 2022 · 2022
Closest in time.
Human Help Wanted: Why AI Is Terrible at Content Moderation
Ben Dickson. 2019 · 2022
Closest in time.
CrowdWorkSheets: Accounting for Individual and Collective Identities Underlying Crowdsourced Dataset Annotation. In 2022 ACM Conference on Fairness, Accountability, and Transparency . ACM, Seoul, 2342–2351
Mark Díaz, Ian Kivlichan, Rachel Rosen, Dylan Baker, Razvan Amironesei, Vinodkumar Prabhakaran, and Emily Denton. 2022 · 2022
Closest in time.
Should attention be all we need? The epistemic and ethical implications of unification in machine learning. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency
Nic Fishman and Leif Hancox-Li. 2022 · 2022
Closest in time.
Evaluation Gaps in Machine Learning Practice. In 2022 ACM Conference on Fairness, Accountability, and Transparency . ACM, Seoul, 1859–1876
Ben Hutchinson, Negar Rostamzadeh, Christina Greer, Katherine Heller, and Vinodkumar Prabhakaran. 2022 · 2022
Closest in time.
A Culture of Ethical AI: Report
Ada Lovelace Institute, CIFAR, and Partnership on AI. 2022 · 2022
Closest in time.
The Ghost in the Machine has an American accent: value conflict in GPT-3
Rebecca L Johnson, Giada Pistilli, Natalia Menédez-González, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, and Donald Jay Bertulfo. 2022 · 2022
Closest in time.
Metaethical Perspectives on ’Benchmarking’ AI Ethics
Travis LaCroix and Alexandra Sasha Luccioni. 2022 · 2022
Closest in time.
Holistic Evaluation of Language Models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2022 · 2022
Closest in time.
Ethical Review Guidelines
Sasha Luccioni, William Isaac, Cherie Poland, Deborah Raji, Samy Bengio, Kate Crawford, Jeanne Fromer, Iason Gabriel, Amanda Levendowski, and Marc’Aurelio Ranzato. 2021 · 2022
Closest in time.
Disordering Datasets: Sociotechnical Misalignments in AI-Mediated Behavioral Health
Varoon Mathur, Caitlin Lustig, and Elizabeth Kaziunas. 2022 · 2022
Closest in time.
The devolution of eugenic practices: Sexual and reproductive health and oppression of people with intellectual disability
David McConnell and Shanon Phelan. 2022 · 2022
Closest in time.
Image Classification on ImageNet
Papers With Code. 2022a · 2022
Closest in time.
Machine Translation on WMT2014 English-German
Papers With Code. 2022b · 2022
Closest in time.
Object Detection on COCO test-dev
Papers With Code. 2022c · 2022
Closest in time.
Semantic Segmentation on ADE20K
Papers With Code. 2022d · 2022
Closest in time.
Speech Recognition on LibriSpeech test-clean
Papers With Code. 2022e · 2022
Closest in time.
Ableism in Education: Rethinking School Practices and Policies
Gillian Parekh. 2022 · 2022
Closest in time.
Annual Research Review: Shifting from ‘normal science’to neurodiversity in autism science
Elizabeth Pellicano and Jacquiline den Houting. 2022 · 2022
Closest in time.
Looking before we leap: Ethical review processes for AI and data science research
Mylene Petermann, Niccolo Tempini, Ismael Kherroubi Garcia, Kirstie Whitaker, and Andrew Strait. 2022 · 2022
Closest in time.
Advancing ethics review practices in AI research
Madhulika Srikumar, Rebecca Finlay, Grace Abuhamad, Carolyn Ashurst, Rosie Campbell, Emily Campbell-Ratcliffe, Hudson Hongo, Sara R. Jordan, Joseph Lindley, Aviv Ovadya, and Joelle Pineau. 2022 · 2022
Closest in time.
Announcing the NeurIPS Code of Ethics
Samy Bengio, Sasha Luccioni, and Inioluwa Deborah Raji. 2023 · 2023
Closest in time.
A Critical Analysis of Standardized Testing in Speech and Language Therapy
Vishnu KK Nair, Warda Farah, and Ian Cushing. 2023 · 2023
Closest in time.