Fetching the paper…
Reading the bibliography…
The "double descent" risk curve was proposed to qualitatively describe the out-of-sample prediction accuracy of variably-parameterized machine learning models.
“How many variables should be entered in a regression equation?”
Leo Breiman and David Freedman · 1983
Earlier work this paper cites.
“Linear and nonlinear extension of the pseudo-inverse solution for learning boolean functions”
F Vallet, J-G Cailton and Ph Refregier · 1989
Earlier work this paper cites.
“On the ability of the optimal perceptron to generalise”
M Opper, W Kinzel, J Kleinz and R Nehl · 1990
Earlier work this paper cites.
“Eigenvalues of covariance matrices: Application to neural-network learning”
Yann Le, Ido Kanter and Sara Solla · 1991
Earlier work this paper cites.
“Generalization in a linear perceptron in the presence of noise”
Anders Krogh and John Hertz · 1992
Earlier work this paper cites.
“The statistical mechanics of learning a rule”
Timothy Watkin, Albrecht Rau and Michael Biehl · 1993
Earlier work this paper cites.
“Dynamics of batch training in a perceptron”
Siegfried Bös and Manfred Opper · 1998
Earlier work this paper cites.
“Learning probability distributions”, 2000
Sanjoy Dasgupta · 2000
Earlier work this paper cites.
“An elementary proof of a theorem of Johnson and Lindenstrauss”
Sanjoy Dasgupta and Anupam Gupta · 2003
Cited alongside, same era.
“Random features for large-scale kernel machines”
Ali Rahimi and Benjamin Recht · 2008
Cited alongside, same era.
“Smallest singular value of a random rectangular matrix”
Mark Rudelson and Roman Vershynin · 2009
Cited alongside, same era.
“Limiting empirical singular value distribution of restrictions of discrete Fourier transform matrices”
Brendan Farrell · 2011
Cited alongside, same era.
“Concentration inequalities for sampling without replacement”
Rémi Bardenet and Odalric-Ambrym Maillard · 2015
Cited alongside, same era.
“High-dimensional dynamics of generalization error in neural networks”
“High-dimensional probability: An introduction with applications in data science”
Roman Vershynin · 2018
Later among the works it cites.
“Reconciling modern machine learning practice and the bias-variance trade-off”
Mikhail Belkin, Daniel Hsu, Siyuan Ma and Soumik Mandal · 2019
Closest in time.
“Two models of double descent for weak features”
Mikhail Belkin, Daniel Hsu and Ji Xu · 2019
Closest in time.
“Surprises in High-Dimensional Ridgeless Least Squares Interpolation”
Trevor Hastie, Andrea Montanari, Saharon Rosset and Ryan Tibshirani · 2019
Closest in time.
“On the number of variables to use in principal component regression”
Ji Xu and Daniel Hsu · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Madhu Advani and Andrew Saxe · 2017
Cited alongside, same era.
“A Modern Take on the Bias-Variance Tradeoff in Neural Networks”
Brady Neal, Sarthak Mittal, Aristide Baratin, Vinayak Tantia, Matthew Scicluna, Simon Lacoste-Julien and Ioannis Mitliagkas · 2018
Cited alongside, same era.
“A jamming transition from under-to over-parametrization affects loss landscape and generalization”
Stefano Spigler, Mario Geiger, Stéphane d’Ascoli, Levent Sagun, Giulio Biroli and Matthieu Wyart · 2018
Cited alongside, same era.
Peter Bartlett, Philip Long, Gábor Lugosi and Alexander Tsigler · 2020
Closest in time.
“A brief prehistory of double descent”
Marco Loog, Tom Viering, Alexander Mey, Jesse Krijthe and David Tax · 2020
Closest in time.
“Harmless interpolation of noisy data in regression”
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian and Anant Sahai · 2020
Closest in time.