Fetching the paper…
Reading the bibliography…
Deep neural networks are known to exhibit a `double descent' behavior as the number of parameters increases.
Distribution of eigenvalues for some sets of random matrices
Vladimir A Marčenko and Leonid Andreevich Pastur · 1967
Earlier work this paper cites.
Linear and nonlinear extension of the pseudo-inverse solution for learning boolean functions
F Vallet, J-G Cailton, and Ph Refregier · 1989
Earlier work this paper cites.
On the ability of the optimal perceptron to generalise
M Opper, W Kinzel, J Kleinz, and R Nehl · 1990
Earlier work this paper cites.
Eigenvalues of covariance matrices: Application to neural-network learning
Yann Le Cun, Ido Kanter, and Sara A Solla · 1991
Earlier work this paper cites.
Generalization in a linear perceptron in the presence of noise
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
The statistical mechanics of learning a rule
Timothy LH Watkin, Albrecht Rau, and Michael Biehl · 1993
Earlier work this paper cites.
Statistical mechanics of learning: Generalization
Manfred Opper · 1995
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Cited alongside, same era.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
More data can hurt for linear regression: Sample-wise double descent
Preetum Nakkiran · 2019
Later among the works it cites.
Sgd on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Cited alongside, same era.
Insights on representational similarity in neural networks with canonical correlation
Ari Morcos, Maithra Raghu, and Samy Bengio · 2018
Cited alongside, same era.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Cited alongside, same era.
Later among the works it cites.
A brief prehistory of double descent
Marco Loog, Tom Viering, Alexander Mey, Jesse H Krijthe, and David MJ Tax · 2020
Later among the works it cites.
Reply to loog et al.: Looking beyond the peaking phenomenon
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2020
Later among the works it cites.
Early stopping in deep networks: Double descent and how to eliminate it
Reinhard Heckel and Fatih Furkan Yilmaz · 2020
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Curtis G Northcutt, Anish Athalye, and Jonas Mueller · 2021
Closest in time.
On the geometry of generalization and memorization in deep neural networks
Cory Stephenson, Suchismita Padhy, Abhinav Ganesh, Yue Hui, Hanlin Tang, and SueYeon Chung · 2021
Closest in time.