Fetching the paper…
Reading the bibliography…
Double descent is a surprising phenomenon in machine learning, in which as the number of model parameters grows relative to the number of data, test error drops as models grow ever larger into the highly overparameterized (data undersampled) regime.
Statistical mechanics of learning: Generalization
Manfred Opper · 1995
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
A jamming transition from under-to over-parametrization affects loss landscape and generalization
Stefano Spigler, Mario Geiger, Stéphane d’Ascoli, Levent Sagun, Giulio Biroli, and Matthieu Wyart · 2018
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
Double descent in the condition number
Tomaso Poggio, Gil Kur, and Andrzej Banburski · 2019
Earlier work this paper cites.
Understanding double descent requires a fine-grained bias-variance decomposition
Ben Adlam and Jeffrey Pennington · 2020
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani, Andrew M Saxe, and Haim Sompolinsky · 2020
Cited alongside, same era.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Cited alongside, same era.
Optimal regularization can mitigate double descent
Preetum Nakkiran, Prayaag Venkat, Sham Kakade, and Tengyu Ma · 2020
Cited alongside, same era.
Multiple descent: Design your own generalization curve
Lin Chen, Yifei Min, Mikhail Belkin, and Amin Karbasi · 2021
Cited alongside, same era.
Deep double descent: Where bigger models and more data hurt
The geometry of over-parameterized regression and adversarial perturbations
Jason W Rocks and Pankaj Mehta · 2021
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2022
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Later among the works it cites.
Bias-variance decomposition of overparameterized regression with random linear features
Jason W Rocks and Pankaj Mehta · 2022
Later among the works it cites.
Memorizing without overfitting: Bias, variance, and interpolation in overparameterized models
Jason W Rocks and Pankaj Mehta · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Cited alongside, same era.
Double descent in the condition number
Tom Henighan, Shan Carter, Tristan Hume, Nelson Elhage, Robert Lasenby, Stanislav Fort, Nicholas Schiefer, and Christopher Olah · 2023
Closest in time.