Fetching the paper…
Reading the bibliography…
This graduate textbook on machine learning tells a story of how patterns in data support predictions and consequential actions.
Sur les applications de la théorie des probabilités aux experiences agricoles: Essai des principes
Jerzy Neyman · 1923
Earlier work this paper cites.
On the use and interpretation of certain test criteria for purposes of statistical inference: Part I
Jerzy Neyman and Egon S. Pearson · 1928
Earlier work this paper cites.
On the problem of the most efficient tests of statistical hypotheses
Jerzy Neyman and Egon S. Pearson · 1933
Earlier work this paper cites.
Contributions to the theory of statistical estimation and testing hypotheses
Abraham Wald · 1939
Earlier work this paper cites.
Functions aleatoire de second ordre
Michel Loève · 1946
Earlier work this paper cites.
Über lineare Methoden in der Wahrscheinlichkeitsrechnung
Kari Karhunen · 1947
Earlier work this paper cites.
Theory of reproducing kernels
N. Aronszajn · 1950
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
The interpretation of interaction in contingency tables
Edward H. Simpson · 1951
Earlier work this paper cites.
The theory of signal detectability
W. Wesley Peterson, Theodore G. Birdsall, and William C. Fox · 1954
Earlier work this paper cites.
A decision-making theory of visual detection
Wilson P. Tanner Jr. and John A. Swets · 1954
Earlier work this paper cites.
Dynamic programming under uncertainty with a quadratic criterion function
Herbert A. Simon · 1956
Earlier work this paper cites.
An optimum character recognition system using decision functions
Chao Kong Chow · 1957
Earlier work this paper cites.
A note on certainty equivalence in dynamic planning
Henri Theil · 1957
Earlier work this paper cites.
The New York Times
New navy device learns by doing; psychologist shows embryo of computer designed to read and grow wiser · 1958
Earlier work this paper cites.
The perceptron: A probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
Two theorems of statistical separability in the perceptron
Frank Rosenblatt · 1958
Earlier work this paper cites.
A generalized scanner for pattern- and character-recognition studies
Wilbur H. Highleyman and Louis A. Kamentsky · 1959
Earlier work this paper cites.
Adaptive switching circuits
Bernard Widrow and Marcian E. Hoff · 1960
Earlier work this paper cites.
Comments on a character recognition method of bledsoe and browning
Wilbur H. Highleyman and Louis A. Kamentsky · 1960
Earlier work this paper cites.
Dual control theory
Aleksandr Aronovich Feldbaum · 1960
Earlier work this paper cites.
Some mathematical models of learning
Seymour A. Papert · 1961
Earlier work this paper cites.
An approach to time series analysis
Emmanuel Parzen · 1961
Earlier work this paper cites.
Character recognition system, 1961
Wilbur H. Highleyman · 1961
Earlier work this paper cites.
Further results on the n-tuple pattern recognition method
Woodrow Wilson Bledsoe · 1961
Earlier work this paper cites.
Linear decision functions, with application to pattern recognition
Wilbur H. Highleyman · 1962
Earlier work this paper cites.
On convergence proofs on perceptrons
Albert B. J. Novikoff · 1962
Earlier work this paper cites.
Principles of neurodynamics: Perceptions and the theory of brain mechanisms
Frank Rosenblatt · 1962
Earlier work this paper cites.
The perceptron: A model for brain functioning
Hans-Dieter Block · 1962
Earlier work this paper cites.
A recognition method using neighbor dependence
Chao Kong Chow · 1962
Earlier work this paper cites.
The design and analysis of pattern recognition experiments
Wilbur H. Highleyman · 1962
Earlier work this paper cites.
Data for character recognition studies
Wilbur H. Highleyman · 1963
Earlier work this paper cites.
About convergence of random search method in extremal control of multi-parameter systems
Leonard A. Rastrigin · 1963
Earlier work this paper cites.
Linear and nonlinear separation of patterns by linear programming
Olvi L. Mangasarian · 1965
Earlier work this paper cites.
The Robbins-Monro process and the method of potential functions
M. A. Aizerman, E. M. Braverman, and L. I. Rozonoer · 1965
Earlier work this paper cites.
Experiments with highleyman’s data
John H. Munson, Richard O. Duda, and Peter E. Hart · 1968
Earlier work this paper cites.
Pattern classification and scene analysis
Richard O. Duda, Peter E. Hart, and David G. Stork · 1973
Earlier work this paper cites.
Theory of Pattern Recognition: Statistical Learning Problems
Vladimir Vapnik and Alexey Chervonenkis · 1974
Earlier work this paper cites.
Sex bias in graduate admissions: Data from Berkeley
Peter J. Bickel, Eugene A. Hammel, and J. William O’Connell · 1975
Earlier work this paper cites.
Evolutionsstrategie und numerische Optimierung
Hans-Paul Schwefel · 1975
Earlier work this paper cites.
Applications of control theory to macroeconomics
David Kendrick · 1976
Earlier work this paper cites.
Approximations of dynamic programs, I
Ward Whitt · 1978
Earlier work this paper cites.
The history of statistics in the 17th and 18th centuries against the changing background of intellectual, scientific and religious thought
Karl Pearson and Egon S. Pearson · 1981
Earlier work this paper cites.
Remarques sur un résultat non publié de B. Maurey
Gilles Pisier · 1981
Earlier work this paper cites.
Machine learning: An artificial intelligence approach
Ryszard S. Michalski, Jamie G. Carbonell, and Tom M. Mitchell, editors · 1983
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. Nemirovski and D. Yudin · 1983
Earlier work this paper cites.
A note on screening regression equations
David A. Freedman · 1983
Earlier work this paper cites.
N-widths in approximation theory
Allan Pinkus · 1985
Earlier work this paper cites.
The control revolution: Technological and economic origins of the information society
James Beniger · 1986
Earlier work this paper cites.
Parallel distributed processing
James L. McClelland, David E. Rumelhart, and PDP Research Group · 1986
Earlier work this paper cites.
The complexity of Markov Decision Processes
Christos H. Papadimitriou and John N. Tsitsiklis · 1987
Earlier work this paper cites.
What size net gives valid generalization?
Eric B. Baum and David Haussler · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Clinical versus actuarial judgment
Robyn M. Dawes, David Faust, and Paul E. Meehl · 1989
Earlier work this paper cites.
Spline Models for Observational Data
Grace Wahba · 1990
Earlier work this paper cites.
The cybernetics group
Steve J. Heims · 1991
Earlier work this paper cites.
Neural-network and k-nearest-neighbor classifiers
Jane Bromley and Eduard Sackinger · 1991
Earlier work this paper cites.
Statistical models and shoe leather
David A. Freedman · 1991
Earlier work this paper cites.
A simple lemma on greedy approximation in Hilbert space and convergence rates for projection pursuit regression and neural network training
Lee K. Jones · 1992
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
James C. Spall · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R. Barron · 1993
Earlier work this paper cites.
Hinging hyperplanes for regression, classification, and function approximation
Leo Breiman · 1993
Earlier work this paper cites.
DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1
John S. Garofolo, Lori F. Lamel, William M. Fisher, Jonathan G. Fiscus, and David S. Pallett · 1993
Earlier work this paper cites.
Toward efficient agnostic learning
Michael J. Kearns, Robert E. Schapire, and Linda M. Sellie · 1994
Earlier work this paper cites.
Interior-point polynomial methods in convex programming
Yurii Nesterov and Arkadi Nemirovskii · 1994
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Online Q-learning using connectionist systems
Gavin Adrian Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
An upper bound on the loss from approximate optimal-value functions
Satinder P. Singh and Richard C. Yee · 1994
Earlier work this paper cites.
NIST special database 19 handprinted forms and characters database
Patrick J. Grother · 1995
Earlier work this paper cites.
Adaptive dual control methods: An overview
Björn Wittenmark · 1995
Earlier work this paper cites.
Exponentially many local minima for single neurons
Peter Auer, Mark Herbster, and Manfred K. Warmuth · 1996
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Signals and Systems
Alan V. Oppenheim, Alan S. Willsky, and S. Hamid Nawab · 1997
Earlier work this paper cites.
Statistical Larning Theory
Vladimir Vapnik · 1998
Earlier work this paper cites.
A tutorial on support vector machines for pattern recognition
Christopher J. C. Burges · 1998
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
Robert E Schapire, Yoav Freund, Peter Bartlett, Wee Sun Lee, et al · 1998
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L. Bartlett · 1998
Earlier work this paper cites.
On the infeasibility of training neural networks with small mean-squared error
Van H. Vu · 1998
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L. Bartlett · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
The real world of technology
Ursula Franklin · 1999
Earlier work this paper cites.
Causation, Prediction, and Search
Peter Spirtes, Clark N. Glymour, Richard Scheines, David Heckerman, Christopher Meek, Gregory Cooper, and Thomas Richardson · 2000
Earlier work this paper cites.
A comparison of observational studies and randomized, controlled trials
Kjell Benson and Arthur J. Hartz · 2000
Earlier work this paper cites.
Randomized, controlled trials, observational studies, and the hierarchy of research designs
John Concato, Nirav Shah, and Ralph I. Horwitz · 2000
Earlier work this paper cites.
A survey of computational complexity results in systems and control
Vincent D. Blondel and John N. Tsitsiklis · 2000
Cited alongside, same era.
Fragile families: Sample and design
Nancy E. Reichman, Julien O. Teitler, Irwin Garfinkel, and Sara S. McLanahan · 2001
Cited alongside, same era.
Between human and machine: feedback, control, and computing before cybernetics
David A. Mindell · 2002
Cited alongside, same era.
Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond
Bernhard Schölkopf and Alexander J. Smola · 2002
Cited alongside, same era.
Training invariant support vector machines
Dennis Decoste and Bernhard Schölkopf · 2002
Cited alongside, same era.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Cited alongside, same era.
Inherent trade-offs in the fair determination of risk scores
Jon M. Kleinberg, Sendhil Mullainathan, and Manish Raghavan · 2017
Later among the works it cites.
Fair prediction with disparate impact: A study of bias in recidivism prediction instruments
Alexandra Chouldechova · 2017
Later among the works it cites.
Perceptrons: An introduction to computational geometry
Marvin Minsky and Seymour A. Papert · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2002
Cited alongside, same era.
Empirical margin distributions and bounding the generalization error of combined classifiers
Vladimir Koltchinskii and Dmitry Panchenko · 2002
Cited alongside, same era.
Postmenopausal Hormone Replacement Therapy and the Primary Prevention of Cardiovascular Disease
Linda L. Humphrey, Benjamin K.S. Chan, and Harold C. Sox · 2002
Cited alongside, same era.
Evolution Strategies—a comprehensive introduction
Hans-Georg Beyer and Hans-Paul Schwefel · 2002
Cited alongside, same era.
Respect the unstable
Gunter Stein · 2003
Cited alongside, same era.
Kernel methods for pattern analysis
John Shawe-Taylor and Nello Cristianini · 2004
Cited alongside, same era.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Later among the works it cites.
Theory of deep learning III: Generalization properties of SGD
Chiyuan Zhang, Qianli Liao, Alexander Rakhlin, Karthik Sridharan, Brando Miranda, Noah Golowich, and Tomaso Poggio · 2017
Later among the works it cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2017
Later among the works it cites.
Spectrally-normalized margin bounds for neural networks
Peter L. Bartlett, Dylan J. Foster, and Matus J. Telgarsky · 2017
Later among the works it cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction (Corrected 12th printing)
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2017
Later among the works it cites.
What led computer vision to deep learning?
Jitendra Malik · 2017
Later among the works it cites.
Exposed! A survey of attacks on private data
Cynthia Dwork, Adam Smith, Thomas Steinke, and Jonathan Ullman · 2017
Later among the works it cites.
EFF AI progress measurement project
Peter Eckersley, Yomna Nasser, et al · 2017
Later among the works it cites.
Elements of Causal Inference
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2017
Later among the works it cites.
The state of applied econometrics: Causality and policy evaluation
Susan Athey and Guido W. Imbens · 2017
Later among the works it cites.
Dynamic Programming and Optimal Control
Dimitri P. Bertsekas · 2017
Later among the works it cites.
Predictive Control for Linear and Hybrid Systems
Francesco Borrelli, Alberto Bemporad, and Manfred Morari · 2017
Later among the works it cites.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2017
Later among the works it cites.
Emerging trends: A tribute to Charles Wayne
Kenneth Ward Church · 2018
Later among the works it cites.
Artificial unintelligence: How computers misunderstand the world
Meredith Broussard · 2018
Later among the works it cites.
Automating inequality: How high-tech tools profile, police, and punish the poor
Virginia Eubanks · 2018
Later among the works it cites.
Algorithms of oppression: How search engines reinforce racism
Safiya Umoja Noble · 2018
Later among the works it cites.
Measurement theory and applications for the social sciences
Deborah L. Bandalos · 2018
Later among the works it cites.
Learning certifiably optimal rule lists for categorical data
Elaine Angelino, Nicholas Larus-Stone, Daniel Alabi, Margo Seltzer, and Cynthia Rudin · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Later among the works it cites.
Group normalization
Yuxin Wu and Kaiming He · 2018
Later among the works it cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Later among the works it cites.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2018
Later among the works it cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Later among the works it cites.
How copyright law can fix artificial intelligence’s implicit bias problem
Amanda Levendowski · 2018
Later among the works it cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2018
Later among the works it cites.
The Book of Why: The New Science of Cause and Effect
Judea Pearl and Dana Mackenzie · 2018
Later among the works it cites.
Understanding and misunderstanding randomized controlled trials
Angus Deaton and Nancy Cartwright · 2018
Later among the works it cites.
Estimation and inference of heterogeneous treatment effects using random forests
Stefan Wager and Susan Athey · 2018
Later among the works it cites.
Quasi-experimental causality in neuroscience and behavioural research
Ioana E. Marinescu, Patrick N. Lawlor, and Konrad P. Kording · 2018
Later among the works it cites.
A smoothed analysis of the greedy algorithm for the linear contextual bandit problem
Sampath Kannan, Jamie H. Morgenstern, Aaron Roth, Bo Waggoner, and Zhiwei Steven Wu · 2018
Later among the works it cites.
Alberto Bietti, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Abstract Dynamic Programming
Dimitri P. Bertsekas · 2018
Later among the works it cites.
Artificial intelligence—the revolution hasn’t happened yet
Michael I. Jordan · 2019
Later among the works it cites.
Race after Technology
Ruha Benjamin · 2019
Later among the works it cites.
50 years of test (un) fairness: Lessons for machine learning
Ben Hutchinson and Margaret Mitchell · 2019
Later among the works it cites.
Fairness and Machine Learning
Solon Barocas, Moritz Hardt, and Arvind Narayanan · 2019
Later among the works it cites.
Machine, learning, 1951
Jef Akst · 2019
Later among the works it cites.
Why random reshuffling beats stochastic gradient descent
Mert Gürbüzbalaban, Asu Ozdaglar, and Pablo A Parrilo · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, and Zhifeng Chen · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
Vocal features: From voice identification to speech recognition by machine
Xiaochang Li and Mara Mills · 2019
Later among the works it cites.
Cold case: The lost MNIST digits
Chhavi Yadav and Léon Bottou · 2019
Later among the works it cites.
Ghost work: how to stop Silicon Valley from building a new global underclass
Mary L. Gray and Siddharth Suri · 2019
Later among the works it cites.
Model similarity mitigates test set overuse
Horia Mania, John Miller, Ludwig Schmidt, Moritz Hardt, and Benjamin Recht · 2019
Later among the works it cites.
Hila Gonen and Yoav Goldberg · 2019
Later among the works it cites.
Do ImageNet classifiers generalize to ImageNet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Later among the works it cites.
Social data: Biases, methodological pitfalls, and ethical boundaries
Alexandra Olteanu, Carlos Castillo, Fernando Diaz, and Emre Kiciman · 2019
Later among the works it cites.
Data colonialism: Rethinking big data’s relation to the contemporary subject
Nick Couldry and Ulises A. Mejias · 2019
Later among the works it cites.
Causality for machine learning
Bernhard Schölkopf · 2019
Later among the works it cites.
A comparison of approaches to advertising measurement: Evidence from big field experiments at facebook
Brett R. Gordon, Florian Zettelmeyer, Neha Bhargava, and Dan Chapsky · 2019
Later among the works it cites.
How We Cooperate: A Theory of Kantian Optimization
John E. Roemer · 2019
Later among the works it cites.
A tour of reinforcement learning: The view from continuous control
Benjamin Recht · 2019
Later among the works it cites.
Certainty equivalence is efficient for linear quadratic control
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Later among the works it cites.
Reinforcement Learning and Optimal Control
Dimitri P. Bertsekas · 2019
Later among the works it cites.
Human language technology
Mark Liberman and Charles Wayne · 2020
Later among the works it cites.
Neural kernels without tangents
Vaishaal Shankar, Alex Fang, Wenshuo Guo, Sara Fridovich-Keil, Jonathan Ragan-Kelley, Ludwig Schmidt, and Benjamin Recht · 2020
Later among the works it cites.
On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai · 2020
Later among the works it cites.
Compressive sensing with un-trained neural networks: Gradient descent finds a smooth approximation
Reinhard Heckel and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
https://myrtle.ai/learn/how-to-train-your-resnet/ , 2020
David Page · 2020
Later among the works it cites.
Racial disparities in automated speech recognition
Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, Connor Toups, John R Rickford, Dan Jurafsky, and Sharad Goel · 2020
Later among the works it cites.
personal communication, 2020
David Aha · 2020
Later among the works it cites.
Why do classifier accuracies show linear trends under distribution shift?
Horia Mania and Suvrit Sra · 2020
Later among the works it cites.
Evaluating machine accuracy on ImageNet
Vaishaal Shankar, Rebecca Roelofs, Horia Mania, Alex Fang, Benjamin Recht, and Ludwig Schmidt · 2020
Later among the works it cites.
Lessons from archives: Strategies for collecting sociocultural data in machine learning
Eun Seo Jo and Timnit Gebru · 2020
Later among the works it cites.
Data and its (dis)contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton, and Alex Hanna · 2020
Later among the works it cites.
Causal Inference: What If
Miguel A. Hernán and James Robins · 2020
Later among the works it cites.
Subprime Attention Crisis
Tim Hwang · 2020
Later among the works it cites.
Naive exploration is optimal for online LQR
Max Simchowitz and Dylan Foster · 2020
Later among the works it cites.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, Sham M. Kakade, and Wen Sun · 2020
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
Alekh Agarwal, Sham M. Kakade, and Lin F. Yang · 2020
Later among the works it cites.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Optimization for Data Analysis
Stephen J. Wright and Benjamin Recht · 2021
Closest in time.
Interpolating classifiers make few mistakes
Tengyuan Liang and Benjamin Recht · 2021
Closest in time.
Generalization in overparameterized models
Moritz Hardt · 2021
Closest in time.
Tight hardness results for training depth-2 ReLU networks
Surbhi Goel, Adam Klivans, Pasin Manurangsi, and Daniel Reichman · 2021
Closest in time.
Retiring adult: New datasets for fair machine learning
Frances Ding, Moritz Hardt, John Miller, and Ludwig Schmidt · 2021
Closest in time.
http://yann.lecun.com/exdb/mnist/
Yann LeCun · 2021
Closest in time.