Fetching the paper…
Reading the bibliography…
Rapid model validation via the train-test paradigm has been a key driver for the breathtaking progress in machine learning and AI.
“Construct validity in psychological tests”
Lee Cronbach and Paul Meehl · 1955
Earlier work this paper cites.
“On a class of skew distribution functions”
Herbert Simon · 1955
Earlier work this paper cites.
“The structure of scientific revolutions”
Thomas Kuhn · 1970
Earlier work this paper cites.
“Meaning and Values in Test Validation: The Science and Ethics of Assessment”
Samuel Messick · 1989
Earlier work this paper cites.
“On the Epistemic Limits of Personalized Prediction”
Lucas Monteiro, Carol Long, Berk Ustun and Flavio Calmon · 1991
Earlier work this paper cites.
“Birthday paradox, coupon collectors, caching algorithms and self-organizing search”
Philippe Flajolet, Danièle Gardy and Loÿs Thimonier · 1992
Earlier work this paper cites.
“The lack of a priori distinctions between learning algorithms”
David Wolpert · 1996
Earlier work this paper cites.
“Emergence of scaling in random networks”
Albert-László Barabási and Réka Albert · 1999
Earlier work this paper cites.
“Is Science Value Free?: Values and Scientific Understanding”
Hugh Lacey · 2005
Earlier work this paper cites.
“The Large-scale structure of semantic networks: Statistical analyses and a model of semantic growth”
Mark Steyvers and Joshua Tenenbaum · 2005
Earlier work this paper cites.
“Improving recommendation lists through topic diversification”
Cai-Nicolas Ziegler, Sean McNee, Joseph Konstan and Georg Lausen · 2005
Earlier work this paper cites.
“DBpedia: A Nucleus for a Web of Open Data”
Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak and Zachary Ives · 2007
Earlier work this paper cites.
“Yago: a core of semantic knowledge”
Fabian Suchanek, Gjergji Kasneci and Gerhard Weikum · 2007
Earlier work this paper cites.
“The End of Theory: The Data Deluge Makes the Scientific Method Obsolete”
Chris Anderson · 2008
Earlier work this paper cites.
“Freebase: a collaboratively created graph database for structuring human knowledge”
Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge and Jamie Taylor · 2008
Earlier work this paper cites.
“Matrix nearness problems with Bregman divergences”
Inderjit Dhillon and Joel Tropp · 2008
Earlier work this paper cites.
“Collaborative prediction and ranking with non-random missing data”
Benjamin Marlin and Richard Zemel · 2009
Earlier work this paper cites.
“Matrix Completion from Power-Law Distributed Samples”
Raghu Meka, Prateek Jain and Inderjit Dhillon · 2009
Earlier work this paper cites.
“Collaborative Filtering in a Non-Uniform World: Learning with the Weighted Trace Norm”
Nathan Srebro and Russ Salakhutdinov · 2010
Earlier work this paper cites.
“Training and testing of recommender systems on data missing not at random”
Harald Steck · 2010
Earlier work this paper cites.
“Popularity versus similarity in growing networks”
Fragkiskos Papadopoulos, Maksim Kitsak, MÁngeles Serrano, Marián Boguná and Dmitri Krioukov · 2012
Earlier work this paper cites.
“Validating the interpretations and uses of test scores”
Michael Kane · 2013
Earlier work this paper cites.
“Understanding Machine Learning: From Theory to Algorithms”
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
“Wikidata: a free collaborative knowledgebase”
Denny Vrandečić and Markus Krötzsch · 2014
Cited alongside, same era.
“Two big challenges in machine learning”, 2015
Léon Bottou · 2015
Cited alongside, same era.
“The MovieLens Datasets: History and Context”
F Harper and Joseph Konstan · 2015
Cited alongside, same era.
“The algebraic combinatorial approach for low-rank matrix completion”
Franz Király, Louis Theran and Ryota Tomioka · 2015
Cited alongside, same era.
“A characterization of deterministic sampling patterns for low-rank matrix completion”
Daniel Pimentel-Alarcón, Nigel Boston and Robert Nowak · 2016
Cited alongside, same era.
“Socially Responsible AI Algorithms: Issues, Purposes, and Challenges”
Lu Cheng, Kush Varshney and Huan Liu · 2021
Later among the works it cites.
“Towards Out-Of-Distribution Generalization: A Survey”
Jiashuo Liu, Zheyan Shen, Yue He, Xingxuan Zhang, Renzhe Xu, Han Yu and Peng Cui · 2021
Later among the works it cites.
“The no-free-lunch theorems of supervised learning”
Tom Sterkenburg and Peter Grünwald · 2021
Later among the works it cites.
“Algorithms on the Bench: Examining Validity of ML Systems in the Public Sphere”, 2022
Rediet Abebe · 2022
Later among the works it cites.
“Model Multiplicity: Opportunities, Concerns, and Solutions”
Emily Black, Manish Raghavan and Solon Barocas · 2022
Later among the works it cites.
“A validity perspective on evaluating the justified use of data-driven decision-making algorithms”
Amanda Coston, Anna Kawakami, Haiyi Zhu, Ken Holstein and Hoda Heidari · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Recommendations as Treatments: Debiasing Learning and Evaluation”
Tobias Schnabel, Adith Swaminathan, Ashudeep Singh, Navin Chandak and Thorsten Joachims · 2016
Cited alongside, same era.
“JAX: composable transformations of Python+NumPy programs”, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne and Qiao Zhang · 2018
Cited alongside, same era.
“How algorithmic confounding in recommendation systems increases homogeneity and decreases utility”
Allison Chaney, Brandon Stewart and Barbara Engelhardt · 2018
Cited alongside, same era.
“Matrix completability analysis via graph k-connectivity”
Dehua Cheng, Natali Ruchansky and Yan Liu · 2018
Cited alongside, same era.
“All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously”
Aaron Fisher, C Rudin and F Dominici · 2018
Cited alongside, same era.
“Statistical Paradises And Paradoxes In Big Data (I) Law Of Large Populations, Big Data Paradox, And The 2016 US Presidential Election”
Xiao-Li Meng · 2018
Cited alongside, same era.
Later among the works it cites.
“Underspecification presents challenges for credibility in modern machine learning”
Alexander D’Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai and D Sculley · 2022
Later among the works it cites.
“Survey of Hallucination in Natural Language Generation”
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Wenliang Dai, Andrea Madotto and Pascale Fung · 2022
Later among the works it cites.
“Breaking Feedback Loops in Recommender Systems with Causal Inference”
Karl Krauth, Yixin Wang and Michael Jordan · 2022
Later among the works it cites.
“There’s more to data than distributions”, 2022
Deborah Raji · 2022
Later among the works it cites.
“Machine Learning has a validity problem”, 2022
Benjamin Recht · 2022
Later among the works it cites.
“On the existence of simpler machine learning models”
Lesia Semenova, Cynthia Rudin and Ronald Parr · 2022
Later among the works it cites.
“Fairness and Machine Learning: Limitations and Opportunities”
Solon Barocas, Moritz Hardt and Arvind Narayanan · 2023
Later among the works it cites.
“Borges and AI”
Léon Bottou and Bernhard Schölkopf · 2023
Later among the works it cites.
“Group fairness without demographics using social networks”
David Liu, Virginie Do, Nicolas Usunier and Maximilian Nickel · 2023
Later among the works it cites.
“GPT-4 Technical Report”
OpenAI et al · 2023
Later among the works it cites.
“Are Emergent Abilities of Large Language Models a Mirage?”
Rylan Schaeffer, Brando Miranda and Sanmi Koyejo · 2023
Later among the works it cites.
“Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models”
Aarohi Srivastava et al · 2023
Later among the works it cites.
“Llama 2: Open Foundation and Fine-Tuned Chat Models”
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Smith, Ranjan Subramanian, Xiaoqing Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov and Thomas Scialom · 2023
Later among the works it cites.
“Predictive multiplicity in probabilistic classification”
Jamelle Watson-Daniels, David Parkes and Berk Ustun · 2023
Later among the works it cites.
“The Llama 3 herd of models”
Abhimanyu Dubey et al · 2024
Closest in time.