Fetching the paper…
Reading the bibliography…
Recent arguments that machine learning (ML) is facing a reproducibility and replication crisis suggest that some published claims in ML research cannot be taken at face value.
Optimizer benchmarking needs to account for hyperparameter tuning
Prabhu Teja Sivaprasad, Florian Mai, Thijs Vogels, Martin Jaggi, and François Fleuret. 2019 · 1910
Earlier work this paper cites.
The test of significance in psychological research
David Bakan. 1966 · 1966
Earlier work this paper cites.
Theory-testing in psychology and physics: A methodological paradox
Paul E Meehl. 1967 · 1967
Earlier work this paper cites.
The language-as-fixed-effect fallacy: A critique of language statistics in psychological research
Herbert H Clark. 1973 · 1973
Earlier work this paper cites.
On the external validity of experiments in consumer research
John G Lynch Jr. 1982 · 1982
Earlier work this paper cites.
A theory of the learnable
Leslie G Valiant. 1984 · 1984
Earlier work this paper cites.
The Likelihood Principle. IMS
James O Berger and Robert L Wolpert. 1988 · 1988
Earlier work this paper cites.
Why summaries of research on psychological theories are often uninterpretable
Paul E Meehl. 1990 · 1990
Earlier work this paper cites.
Comparing the predictive powers of alternative multiple regression models
Michael R Hagerty and V Srinivasan. 1991 · 1991
Earlier work this paper cites.
Statistical power analysis
Jacob Cohen. 1992 · 1992
Earlier work this paper cites.
P values, hypothesis tests, and likelihood: implications for epidemiology of a neglected historical debate
Steven N Goodman. 1993 · 1993
Earlier work this paper cites.
Bayesian Theory
Jose M. Bernardo and Adrian F. M. Smith. 1994 · 1994
Earlier work this paper cites.
Statistics notes: Absence of evidence is not evidence of absence
Douglas G Altman and J Martin Bland. 1995 · 1995
Earlier work this paper cites.
Wavelab and reproducible research
Jonathan B Buckheit and David L Donoho. 1995 · 1995
Earlier work this paper cites.
Automaticity of social behavior: Direct effects of trait construct and stereotype activation on action
John A Bargh, Mark Chen, and Lara Burrows. 1996 · 1996
Earlier work this paper cites.
Learning in the presence of concept drift and hidden contexts
Gerhard Widmer and Miroslav Kubat. 1996 · 1996
Earlier work this paper cites.
Explaining bargaining impasse: The role of self-serving biases
Linda Babcock and George Loewenstein. 1997 · 1997
Earlier work this paper cites.
Statistical Learning Theory
Vladimir Vapnik. 1998 · 1998
Earlier work this paper cites.
Publication bias (the "file-drawer problem") in scientific inference
Jeffrey D Scargle. 1999 · 1999
Earlier work this paper cites.
Stimulus sampling and social psychological experimentation
Gary L Wells and Paul D Windschitl. 1999 · 1999
Earlier work this paper cites.
Statistical modeling: The two cultures
Leo Breiman. 2001 · 2001
Earlier work this paper cites.
Two cheers for P-values?
S Senn. 2001 · 2001
Earlier work this paper cites.
Motivated Reasoning and Performance on the was on Selection Task
Erica Dawson, Thomas Gilovich, and Dennis T Regan. 2002 · 2002
Earlier work this paper cites.
Cultural ways of learning: Individual traits or repertoires of practice
Kris D Gutiérrez and Barbara Rogoff. 2003 · 2003
Earlier work this paper cites.
P values are not error probabilities
Raymond Hubbard and MJ Bayarri. 2003a · 2003
Earlier work this paper cites.
Confusion over measures of evidence (p’s) versus errors ( α \alpha ’s) in classical statistical testing
Raymond Hubbard and María Jesús Bayarri. 2003b · 2003
Earlier work this paper cites.
The Cultural Nature of Human Development
Barbara Rogoff. 2003 · 2003
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Stuart J Russell and Peter Norvig. 2003 · 2003
Earlier work this paper cites.
The estimation of prediction error: Covariance penalties and cross-validation
Bradley Efron. 2004 · 2004
Earlier work this paper cites.
Bayesian inference
Larry Wasserman. 2004 · 2004
Earlier work this paper cites.
Open Graph Benchmark: Datasets for Machine Learning on Graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. 2021 · 2005
Earlier work this paper cites.
Machine learning as an experimental science (revisited). In AAAI Workshop on Evaluation Methods for Machine Learning . 1–5
Chris Drummond. 2006 · 2006
Earlier work this paper cites.
The difference between “significant” and “not significant” is not itself statistically significant
Andrew Gelman and Hal Stern. 2006 · 2006
Earlier work this paper cites.
Psychology as the science of self-reports and finger movements: Whatever happened to actual behavior?
Roy F Baumeister, Kathleen D Vohs, and David C Funder. 2007 · 2007
Earlier work this paper cites.
A practical solution to the pervasive problems of p values
Eric-Jan Wagenmakers. 2007 · 2007
Earlier work this paper cites.
Mostly Harmless Econometrics
Joshua D Angrist and Jörn-Steffen Pischke. 2008 · 2008
Earlier work this paper cites.
Why most discovered true associations are inflated
John P. A. Ioannidis. 2008 · 2008
Earlier work this paper cites.
P-values are random variables
Duncan J Murdoch, Yu-Ling Tsai, and James Adcock. 2008 · 2008
Earlier work this paper cites.
Dataset shift in machine learning
Joaquin Quiñonero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence. 2008 · 2008
Earlier work this paper cites.
Discriminative learning under covariate shift
Steffen Bickel, Michael Brückner, and Tobias Scheffer. 2009 · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. In CVPR . Ieee, 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Of beauty, sex and power: Too little attention has been paid to the statistical challenges in estimating small effects
Andrew Gelman and David Weakliem. 2009 · 2009
Earlier work this paper cites.
The unreasonable effectiveness of data
Alon Halevy, Peter Norvig, and Fernando Pereira. 2009 · 2009
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction . Vol. 2
Trevor Hastie, Robert Tibshirani, and Jerome H Friedman. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
When training and test sets are different: Characterizing learning transfer
Amos Storkey. 2009 · 2009
Earlier work this paper cites.
“Positive” results increase down the hierarchy of the sciences
Daniele Fanelli. 2010 · 2010
Earlier work this paper cites.
The weirdest people in the world?
Joseph Henrich, Steven J Heine, and Ara Norenzayan. 2010 · 2010
Earlier work this paper cites.
To explain or to predict?
Galit Shmueli. 2010 · 2010
Earlier work this paper cites.
P-value precision and reproducibility
Dennis D Boos and Leonard A Stefanski. 2011 · 2011
Earlier work this paper cites.
False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant
Joseph P Simmons, Leif D Nelson, and Uri Simonsohn. 2011 · 2011
Earlier work this paper cites.
Unbiased look at dataset bias. In Proc. of CVPR . IEEE, 1521–1528
Antonio Torralba and Alexei A Efros. 2011 · 2011
Earlier work this paper cites.
Protecting against evaluation overfitting in empirical reinforcement learning. In ADPRL . IEEE, 120–127
Shimon Whiteson, Brian Tanner, Matthew E Taylor, and Peter Stone. 2011 · 2011
Earlier work this paper cites.
Blind retrospection: Why shark attacks are bad for democracy
Christopher H Achen and Larry M Bartels. 2012 · 2012
Earlier work this paper cites.
On the use of cross-validation for time series predictor evaluation
Christoph Bergmeir and José M Benítez. 2012 · 2012
Earlier work this paper cites.
A few useful things to know about machine learning
Pedro Domingos. 2012 · 2012
Earlier work this paper cites.
The psychology of replication and replication in psychology
Gregory Francis. 2012a · 2012
Earlier work this paper cites.
Publication bias and the failure of replication in experimental psychology
Gregory Francis. 2012b · 2012
Earlier work this paper cites.
Ethics and statistics: Ethics and the statistical use of prior information
Andrew Gelman. 2012 · 2012
Earlier work this paper cites.
Results on differential and dependent measurement error of the exposure and the outcome using signed directed acyclic graphs
Tyler J VanderWeele and Miguel A Hernán. 2012 · 2012
Earlier work this paper cites.
A failed replication draws a scathing personal attack from a psychology professor
Ed Yong. 2012 · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent. 2013 · 2013
Earlier work this paper cites.
Power failure: Why small sample size undermines the reliability of neuroscience
Katherine S Button, John Ioannidis, Claire Mokrysz, Brian A Nosek, Jonathan Flint, Emma SJ Robinson, and Marcus R Munafò. 2013 · 2013
Earlier work this paper cites.
The fluctuating female vote: Politics, religion, and the ovulatory cycle
Kristina M Durante, Ashley Rae, and Vladas Griskevicius. 2013 · 2013
Earlier work this paper cites.
P values and statistical practice
Andrew Gelman. 2013 · 2013
Earlier work this paper cites.
The garden of forking paths: Why multiple comparisons can be a problem, even when there is no “fishing expedition” or “p-hacking” and the research hypothesis was posited ahead of time
Andrew Gelman and Eric Loken. 2013 · 2013
Earlier work this paper cites.
Ten simple rules for reproducible computational research
Geir Kjetil Sandve, Anton Nekrutenko, James Taylor, and Eivind Hovig. 2013 · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli. 2013 · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
Nonnaïveté among Amazon Mechanical Turk workers: Consequences and solutions for behavioral researchers
Jesse Chandler, Pam Mueller, and Gabriele Paolacci. 2014 · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Beyond power calculations: Assessing type S (sign) and type M (magnitude) errors
Andrew Gelman and John B. Carlin. 2014 · 2014
Earlier work this paper cites.
Ethics and statistics: The AAA tranche of subprime science
Andrew Gelman and Eric Loken. 2014a · 2014
Earlier work this paper cites.
The statistical crisis in science data-dependent analysis—a “garden of forking paths”—explains why many statistically significant comparisons don’t hold up
Andrew Gelman and Eric Loken. 2014b · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Bayesian Cognitive Modeling: A Practical Course
Michael D Lee and Eric-Jan Wagenmakers. 2014 · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro. 2014 · 2014
Cited alongside, same era.
Provisioning Reproducible Computational Science
Victoria Stodden and Sheila Miguez. 2014 · 2014
Cited alongside, same era.
Vqa: Visual question answering. In Proc. of ICCV . 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Applying machine learning to facilitate autism diagnostics: pitfalls and promises
Daniel Bone, Matthew S Goodwin, Matthew P Black, Chi-Chun Lee, Kartik Audhkhasi, and Shrikanth Narayanan. 2015 · 2015
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Cited alongside, same era.
How cross-validation can go wrong and what to do about it
Marcel Neunhoeffer and Sebastian Sternberg. 2019 · 2019
Later among the works it cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, and Jasper Snoek. 2019 · 2019
Later among the works it cites.
Unbiased look at dataset bias. ICML
B Recht, R Roelofs, L Schmidt, and V Shankar. 2019 · 2019
Later among the works it cites.
Research in social psychology changed between 2011 and 2016: Larger sample sizes, more self-report measures, and more online studies
Kai Sassenberg and Lara Ditrich. 2019 · 2019
Later among the works it cites.
A causal replication framework for designing and assessing replication efforts
Peter M Steiner, Vivian C Wong, and Kylie Anglin. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The loss surfaces of multilayer networks. In Artificial Intelligence and Statistics . PMLR, 192–204
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun. 2015 · 2015
Cited alongside, same era.
Using prediction markets to estimate the reproducibility of scientific research
Anna Dreber, Thomas Pfeiffer, Johan Almenberg, Siri Isaksson, Brad Wilson, Yiling Chen, Brian A. Nosek, and Magnus Johannesson. 2015 · 2015
Cited alongside, same era.
The connection between varying treatment effects and the crisis of unreplicable research: A Bayesian perspective
Andrew Gelman. 2015 · 2015
Cited alongside, same era.
Surrogate science: The idol of a universal method for scientific inference
Gerd Gigerenzer and Julian N Marewski. 2015 · 2015
Cited alongside, same era.
Unequal representation and gender stereotypes in image search results for occupations. In Proc. of the 33rd Annual ACM Conference on Human Factors in Computing Systems . 3819–3828
Matthew Kay, Cynthia Matuszek, and Sean A Munson. 2015 · 2015
Cited alongside, same era.
Estimating the reproducibility of psychological science
Brian A. Nosek et al · 2015
Cited alongside, same era.
Big data’s disparate impact
Solon Barocas and Andrew D Selbst. 2016 · 2016
Cited alongside, same era.
Chhavi Yadav and Léon Bottou. 2019 · 2019
Later among the works it cites.
What matters for on-policy deep actor-critic methods? A large-scale study. In ICLR
Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk, Manu Orsini, Sertan Girgin, Raphaël Marinier, Leonard Hussenot, Matthieu Geist, Olivier Pietquin, Marcin Michalski, Sylvain Gelly, and Olivier Bachem. 2020 · 2020
Later among the works it cites.
Cats are not fish: Deep learning testing calls for out-of-distribution awareness. In Proc. of ASE . 1041–1052
David Berend, Xiaofei Xie, Lei Ma, Lingjun Zhou, Yang Liu, Chi Xu, and Jianjun Zhao. 2020 · 2020
Later among the works it cites.
Survey of machine-learning experimental methods at NeurIPS2019 and ICLR2020
Xavier Bouthillier and Gaël Varoquaux. 2020 · 2020
Later among the works it cites.
With little power comes great responsibility
Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. 2020 · 2020
Later among the works it cites.
Targeting learning: Robust statistics for reproducible research
Jeremy R Coyle, Nima S Hejazi, Ivana Malenica, Rachael V Phillips, Benjamin F Arnold, Andrew Mertens, Jade Benjamin-Chung, Weixin Cai, Sonali Dayal, John M Colford Jr, Alan E Hubbard, and Mark J van der Laan. 2020 · 2020
Later among the works it cites.
Underspecification presents challenges for credibility in modern machine learning
Alexander D’Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D Hoffman, et al · 2020
Later among the works it cites.
The case for formal methodology in scientific reform
Berna Devezer, Danielle J Navarro, Joachim Vandekerckhove, and Erkan Ozge Buzbas. 2020 · 2020
Later among the works it cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2020
Later among the works it cites.
Prediction, estimation, and attribution
Bradley Efron. 2020 · 2020
Later among the works it cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020 · 2020
Later among the works it cites.
Inductive biases for deep learning of higher-level cognition
Anirudh Goyal and Yoshua Bengio. 2020 · 2020
Later among the works it cites.
Don’t stop pretraining: adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020 · 2020
Later among the works it cites.
Transparency and reproducibility in artificial intelligence
Benjamin Haibe-Kains, George Alexandru Adam, Ahmed Hosny, Farnoosh Khodakarami, Levi Waldron, Bo Wang, Chris McIntosh, Anna Goldenberg, Anshul Kundaje, Casey S Greene, et al · 2020
Later among the works it cites.
AI is wrestling with a replication crisis
Will Douglas Heaven. 2020 · 2020
Later among the works it cites.
I tried a bunch of things: The dangers of unexpected overfitting in classification of brain data
Mahan Hosseini, Michael Powell, John Collins, Chloe Callahan-Flintoft, William Jones, Howard Bowman, and Brad Wyble. 2020 · 2020
Later among the works it cites.
Lessons from archives: Strategies for collecting sociocultural data in machine learning. In Proc. of FAccT . 306–316
Eun Seo Jo and Timnit Gebru. 2020 · 2020
Later among the works it cites.
In a forward direction: Analyzing distribution shifts in machine translation test sets over time
Thomas Liao, Benjamin Recht, and Ludwig Schmidt. 2020 · 2020
Later among the works it cites.
AI Adoption in the Enterprise
Roger Magoulas and Steve Swoyer. 2020 · 2020
Later among the works it cites.
A hierarchy of limitations in machine learning
Momin M Malik. 2020 · 2020
Later among the works it cites.
Paths in strange spaces: A comment on preregistration
Danielle Navarro. 2020 · 2020
Later among the works it cites.
Crud (re) defined
Amy Orben and Daniël Lakens. 2020 · 2020
Later among the works it cites.
The sceptical Bayes factor for the assessment of replication success
Samuel Pawel and Leonhard Held. 2020 · 2020
Later among the works it cites.
Performative prediction. In International Conference on Machine Learning . PMLR, 7599–7609
Juan Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, and Moritz Hardt. 2020 · 2020
Later among the works it cites.
Semantic and cognitive tools to aid statistical science: Replace confidence and significance by compatibility and surprise
Zad Rafi and Sander Greenland. 2020 · 2020
Later among the works it cites.
Artificial conscious intelligence
James A Reggia, Garrett E Katz, and Gregory P Davis. 2020 · 2020
Later among the works it cites.
The pitfalls of simplicity bias in neural networks
Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli. 2020 · 2020
Later among the works it cites.
Energy and policy considerations for modern deep learning research. In Proc. of the AAAI Conference on Artificial Intelligence , Vol. 34. 13693–13696
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2020 · 2020
Later among the works it cites.
Is preregistration worthwhile?
Aba Szollosi, David Kellen, Danielle Navarro, Richard Shiffrin, Iris van Rooij, Trisha Van Zandt, and Chris Donkin. 2020 · 2020
Later among the works it cites.
On the value of out-of-distribution testing: An example of Goodhart’s law
Damien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha, Christopher Kanan, and Anton Van Den Hengel. 2020 · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare. 2021 · 2021
Later among the works it cites.
Perspectives on Machine Learning from Psychology’s Reproducibility Crisis
Samuel J Bell and Onno P Kampman. 2021 · 2021
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?. In Proc. of FAccT . 610–623
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Later among the works it cites.
Drawing maps of model space with modular Stan
Ryan Bernstein. 2021 · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Later among the works it cites.
Accounting for variance in machine learning Bbenchmarks. In Machine Learning and Systems (MLSys)
Xavier Bouthillier, Pierre Delaunay, Mirko Bronzi, Assya Trofimov, Brennan Nichyporuk, Justin Szeto, Naz Sepah, Edward Raff, Kanika Madan, Vikram Voleti, Samira Ebrahimi Kahou, Vincent Michalski, Dmitriy Serdyuk, Tal Arbel, Chris Pal, Gaël Varoquaux, and Pascal Vincent. 2021 · 2021
Later among the works it cites.
A troubling analysis of reproducibility and progress in recommender systems research
Maurizio Ferrari Dacrema, Simone Boglio, Paolo Cremonesi, and Dietmar Jannach. 2021 · 2021
Later among the works it cites.
Yehuda Dar, Vidya Muthukumar, and Richard G Baraniuk. 2021 · 2021
Later among the works it cites.
Mostafa Dehghani, Yi Tay, Alexey A Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. 2021 · 2021
Later among the works it cites.
Competency problems: On finding and removing artifacts in language data
Matt Gardner, William Merrill, Jesse Dodge, Matthew E Peters, Alexis Ross, Sameer Singh, and Noah Smith. 2021 · 2021
Later among the works it cites.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021 · 2021
Later among the works it cites.
A loss curvature perspective on training instabilities of deep learning models. In ICLR
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta, Behnam Neyshabur, David Cardoze, George Dahl, Zack Nado, and Orhan Firat. 2021 · 2021
Later among the works it cites.
The disagreement deconvolution: Bringing machine learning performance metrics in line with reality. In Proc. of the 2021 CHI Conference on Human Factors in Computing Systems . 1–14
Mitchell L Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S Bernstein. 2021 · 2021
Later among the works it cites.
Integrating explanation and prediction in computational social science
Jake M Hofman, Duncan J Watts, Susan Athey, Filiz Garip, Thomas L Griffiths, Jon Kleinberg, Helen Margetts, Sendhil Mullainathan, Matthew J Salganik, Simine Vazire, et al · 2021
Later among the works it cites.
Towards accountability for machine learning datasets: Practices from software engineering and infrastructure. In Proc. of FAccT . 560–575
Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell. 2021 · 2021
Later among the works it cites.
Measurement and fairness. In Proc. of FAccT . 375–385
Abigail Z Jacobs and Hanna Wallach. 2021 · 2021
Later among the works it cites.
Are we learning yet? A meta review of evaluation failures across machine learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)
Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt. 2021 · 2021
Later among the works it cites.
Jimmy Lin, Daniel Campos, Nick Craswell, Bhaskar Mitra, and Emine Yilmaz. 2021 · 2021
Later among the works it cites.
Time waits for no one! Analysis and challenges of temporal misalignment
Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, and Noah A Smith. 2021 · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Curtis G Northcutt, Anish Athalye, and Jonas Mueller. 2021 · 2021
Later among the works it cites.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna. 2021 · 2021
Later among the works it cites.
David Picard. 2021 · 2021
Later among the works it cites.
Improving reproducibility in machine learning research: a report from the NeurIPS 2019 reproducibility program
Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, Vincent Larivière, Alina Beygelzimer, Florence d’Alché Buc, Emily Fox, and Hugo Larochelle. 2021 · 2021
Later among the works it cites.
AI and the everything in the whole wide world benchmark
Inioluwa Deborah Raji, Emily M Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. 2021 · 2021
Later among the works it cites.
Putting psychology to the test: Rethinking model evaluation through benchmarking and prediction
Roberta Rocca and Tal Yarkoni. 2021 · 2021
Later among the works it cites.
Do datasets have politics? Disciplinary values in computer vision dataset development
Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton. 2021 · 2021
Later among the works it cites.
Descending through a crowded valley-benchmarking deep learning optimizers. In International Conference on Machine Learning . PMLR, 9367–9376
Robin M Schmidt, Frank Schneider, and Philipp Hennig. 2021 · 2021
Later among the works it cites.
Pre-registration is a game changer. But, like random assignment, it is neither necessary nor sufficient for credible science
Joseph P Simmons, Leif D Nelson, and Uri Simonsohn. 2021 · 2021
Later among the works it cites.
A framework for understanding sources of harm throughout the machine learning life cycle
Harini Suresh and John Guttag. 2021 · 2021
Later among the works it cites.
Data sharing practices and data availability upon request differ across scientific disciplines
Leho Tedersoo, Rainer Küngas, Ester Oras, Kajar Köster, Helen Eenmaa, Äli Leijen, Margus Pedaste, Marju Raju, Anastasiya Astapova, Heli Lukner, et al · 2021
Later among the works it cites.
The piranha problem: Large effects swimming in a small pond
Christopher Tosh, Philip Greengard, Ben Goodrich, Andrew Gelman, Aki Vehtari, and Daniel Hsu. 2021 · 2021
Later among the works it cites.
Misspecification and unreliable interpretations in psychology and social science
Matthew J Vowels. 2021 · 2021
Later among the works it cites.
Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Mike Li, Jong Wook Kim, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt. 2021 · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. 2021 · 2021
Later among the works it cites.
The Dangers of Underclaiming: Reasons for Caution When Reporting How NLP Systems Fail. In Proc. of the 60th Annual Meeting of the ACL (Volume 1: Long Papers) . 7484–7499
Samuel Bowman. 2022 · 2022
Closest in time.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022 · 2022
Closest in time.
We need to think more about how we conduct research
Gerd Gigerenzer. 2022 · 2022
Closest in time.
My recent talk at the NSF town hall focused on the history of the AI winters, how the ML community became "anti-science," and whether the rejection of science will cause a winter for ML theory. I’ll summarize these issues below…
Tom Goldstein. 2022 · 2022
Closest in time.
The generalizability crisis
Tal Yarkoni. 2022 · 2022
Closest in time.
On over-fitting in model selection and subsequent selection bias in performance evaluation
Gavin C Cawley and Nicola L C Talbot. 2010 · 2079
Closest in time.