Fetching the paper…
Reading the bibliography…
A recent line of research has highlighted the existence of a "double descent" phenomenon in deep learning, whereby increasing the number of training examples $N$ causes the generalization error of neural networks to peak when $N$ is of the same order as the number of parameters $P$.
Eigenvalues of covariance matrices: Application to neural-network learning
Yann Le Cun, Ido Kanter, and Sara A Solla · 1991
Earlier work this paper cites.
Generalization in a linear perceptron in the presence of noise
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Small sample size generalization
Robert PW Duin · 1995
Earlier work this paper cites.
Statistical mechanics of generalization
Manfred Opper and Wolfgang Kinzel · 1996
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
Landscape analysis of constraint satisfaction problems
Florent Krzakala and Jorge Kurchan · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
The dipping phenomenon
Marco Loog and Robert PW Duin · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Statistical mechanics of complex neural systems and high dimensional data
Madhu Advani, Subhaneil Lahiri, and Surya Ganguli · 2013
Earlier work this paper cites.
Products of rectangular random matrices: singular values and progressive scattering
Gernot Akemann, Jesper R Ipsen, and Mario Kieburg · 2013
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Spectral density of products of wishart dilute random matrices. part i: the dense case
Thomas Dupic and Isaac Pérez Castillo · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
The simplest model of jamming
Silvio Franz and Giorgio Parisi · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Cited alongside, same era.
Nonlinear random matrix theory for deep learning
Jeffrey Pennington and Pratik Worah · 2017
Cited alongside, same era.
Towards understanding the role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Cited alongside, same era.
Reconciling modern machine learning and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2018
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Later among the works it cites.
A note on the pennington-worah distribution
S Péché et al · 2019
Later among the works it cites.
Eigenvalue distribution of nonlinear models of random matrices
Lucas Benigni and Sandrine Péché · 2019
Later among the works it cites.
A random matrix perspective on mixtures of nonlinearities for deep learning
Ben Adlam, Jake Levinson, and Jeffrey Pennington · 2019
Later among the works it cites.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrew K Lampinen and Surya Ganguli · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
On the spectrum of random features maps of high dimensional data
Zhenyu Liao and Romain Couillet · 2018
Cited alongside, same era.
A jamming transition from under-to over-parametrization affects generalization in deep learning
S Spigler, M Geiger, S d’Ascoli, L Sagun, G Biroli, and M Wyart · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Cited alongside, same era.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2019
Cited alongside, same era.
More data can hurt for linear regression: Sample-wise double descent
Preetum Nakkiran · 2019
Cited alongside, same era.
Preetum Nakkiran, Prayaag Venkat, Sham Kakade, and Tengyu Ma · 2020
Closest in time.
Double trouble in double descent: Bias and variance (s) in the lazy regime
Stéphane d’Ascoli, Maria Refinetti, Giulio Biroli, and Florent Krzakala · 2020
Closest in time.
Generalization error of generalized linear models in high dimensions
Melikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan, and Alyson K Fletcher · 2020
Closest in time.
Benjamin Aubin, Florent Krzakala, Yue M Lu, and Lenka Zdeborová · 2020
Closest in time.
Provable more data hurt in high dimensional least squares estimator
Zeng Li, Chuanlong Xie, and Qinwen Wang · 2020
Closest in time.
Yifei Min, Lin Chen, and Amin Karbasi · 2020
Closest in time.
Multiple descent: Design your own generalization curve
Lin Chen, Yifei Min, Mikhail Belkin, and Amin Karbasi · 2020
Closest in time.
Ben Adlam and Jeffrey Pennington · 2020
Closest in time.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2020
Closest in time.
The gaussian equivalence of generative models for learning with two-layer neural networks
Sebastian Goldt, Galen Reeves, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2020
Closest in time.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew Gordon Wilson and Pavel Izmailov · 2020
Closest in time.