Fetching the paper…
Reading the bibliography…
A lot of Machine Learning (ML) and Deep Learning (DL) research is of an empirical nature.
Teoria statistica delle classi e calcolo delle probabilita
Carlo Bonferroni · 1936
Earlier work this paper cites.
On a test of whether one of two random variables is stochastically larger than the other
Henry B Mann and Donald R Whitney · 1947
Earlier work this paper cites.
Ordered families of distributions
Erich Leo Lehmann · 1955
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Computer intensive methods for hypothesis testing: An introduction
Eric W Noreen · 1989
Earlier work this paper cites.
Individual comparisons by ranking methods
Frank Wilcoxon · 1992
Earlier work this paper cites.
An introduction to the bootstrap
Bradley Efron and Robert J Tibshirani · 1994
Earlier work this paper cites.
Ke-Hai Yuan and Kentaro Hayashi · 2003
Earlier work this paper cites.
Data Structures for Statistical Computing in Python
Wes McKinney · 2010
Earlier work this paper cites.
Evaluating learning algorithms: a classification perspective
Nathalie Japkowicz and Mohak Shah · 2011
Earlier work this paper cites.
An empirical investigation of statistical significance in NLP
Taylor Berg-Kirkpatrick, David Burkett, and Dan Klein · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Bayesian estimation supersedes the t test
John K Kruschke · 2013
Earlier work this paper cites.
Research commentary—too big to fail: large samples and the p-value problem
Mingfeng Lin, Henry C Lucas Jr, and Galit Shmueli · 2013
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Statistical tests, p values, confidence intervals, and power: a guide to misinterpretations
Sander Greenland, Stephen J Senn, Kenneth J Rothman, John B Carlin, Charles Poole, Steven N Goodman, and Douglas G Altman · 2016
Cited alongside, same era.
Models for the assessment of treatment improvement: The ideal and the feasible
PC Álvarez-Esteban, Eustasio del Barrio, Juan Antonio Cuesta-Albertos, and C Matrán · 2017
Cited alongside, same era.
Time for a change: a tutorial for comparing multiple classifiers through bayesian analysis
Alessio Benavoli, Giorgio Corani, Janez Demšar, and Marco Zaffalon · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Not all claims are created equal: Choosing the right statistical approach to assess hypotheses
Erfan Sadeqi Azer, Daniel Khashabi, Ashish Sabharwal, and Dan Roth · 2020
Later among the works it cites.
With little power comes great responsibility
Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky · 2020
Later among the works it cites.
Statistical significance testing for natural language processing
Rotem Dror, Lotem Peled-Cohen, Segev Shlomov, and Roi Reichart · 2020
Later among the works it cites.
Array programming with NumPy
Charles R. Harris, K. Jarrod Millman, Stéfan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fernández del Río, Mark Wiebe, Pearu Peterson, Pierre Gérard-Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, Christoph Gohlke, and Travis E. Oliphant · 2020
Later among the works it cites.
pandas-dev/pandas: Pandas, February 2020
Pandas Development Team · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart · 2018
Cited alongside, same era.
State of the art: Reproducibility in artificial intelligence
Odd Erik Gundersen and Sigbjørn Kjensmo · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Cited alongside, same era.
Model evaluation, model selection, and algorithm selection in machine learning
Sebastian Raschka · 2018
Cited alongside, same era.
Nils Reimers and Iryna Gurevych · 2018
Cited alongside, same era.
Robin M Schmidt, Frank Schneider, and Philipp Hennig · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron Courville, and Marc G Bellemare · 2021
Later among the works it cites.
Accounting for variance in machine learning benchmarks
Xavier Bouthillier, Pierre Delaunay, Mirko Bronzi, Assya Trofimov, Brennan Nichyporuk, Justin Szeto, Nazanin Mohammadi Sepahvand, Edward Raff, Kanika Madan, Vikram Voleti, et al · 2021
Later among the works it cites.
Hyperparameter optimization is deceiving us, and how to stop it
A Feder Cooper, Yucheng Lu, Jessica Forde, and Christopher M De Sa · 2021
Later among the works it cites.
Mostafa Dehghani, Yi Tay, Alexey A Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Bayesian data analysis third edition, 2021
Andrew Gelman, John B Carlin, Hal S Stern, David B Dunson, Aki Vehtari, Donald B Rubin, John Carlin, Hal Stern, Donald Rubin, and David Dunson · 2021
Later among the works it cites.
The role of p-values in judging the strength of evidence and realistic replication expectations
Eric W Gibson · 2021
Later among the works it cites.
Scientific credibility of machine translation research: A meta-evaluation of 769 papers
Benjamin Marie, Atsushi Fujita, and Raphael Rubino · 2021
Later among the works it cites.
Do transformer modifications transfer across implementations and applications?
Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, et al · 2021
Later among the works it cites.
Validity, reliability, and significance
Stefan Riezler and Michael Hagmann · 2021
Later among the works it cites.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam · 2022
Closest in time.