Fetching the paper…
Reading the bibliography…
In this expository note we describe a surprising phenomenon in overparameterized linear regression, where the dimension exceeds the number of samples: there is a regime where the test risk of the estimator found by gradient descent increases with additional samples.
Two models of double descent for weak features
Belkin, M., Hsu, D., and Xu, J. (2019) · 1903
Earlier work this paper cites.
Harmless interpolation of noisy data in regression
Muthukumar, V., Vodrahalli, K., and Sahai, A. (2019) · 1903
Earlier work this paper cites.
A new look at an old problem: A universal learning approach to linear regression
Bibas, K., Fogel, Y., and Feder, M. (2019) · 1905
Earlier work this paper cites.
Benign overfitting in linear regression
Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A. (2019) · 1906
Earlier work this paper cites.
Mitra, P. P. (2019) · 1906
Earlier work this paper cites.
On the risk of minimum-norm interpolants and restricted lower isometry of kernels
Liang, T., Rakhlin, A., and Zhai, X. (2019) · 1908
Earlier work this paper cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S. and Montanari, A. (2019) · 1908
Earlier work this paper cites.
A model of double descent for high-dimensional binary linear classification
Deng, Z., Kammoun, A., and Thrampoulidis, C. (2019) · 1911
Earlier work this paper cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I. (2019) · 1912
Cited alongside, same era.
Distribution of eigenvalues for some sets of random matrices
Marčenko, V. A. and Pastur, L. A. (1967) · 1967
Cited alongside, same era.
Statistical mechanics of learning: Generalization
Opper, M. (1995) · 1995
Cited alongside, same era.
Learning to generalize
Opper, M. (2001) · 2001
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S. and Saxe, A. M. (2017) · 2017
Cited alongside, same era.
Reconciling modern machine learning and the bias-variance trade-off
Just interpolate: Kernel" ridgeless" regression can generalize
Liang, T. and Rakhlin, A. (2018) · 2018
Later among the works it cites.
A modern take on the bias-variance tradeoff in neural networks
Neal, B., Mittal, S., Baratin, A., Tantia, V., Scicluna, M., Lacoste-Julien, S., and Mitliagkas, I. (2018) · 2018
Later among the works it cites.
A jamming transition from under-to over-parametrization affects loss landscape and generalization
Spigler, S., Geiger, M., d’Ascoli, S., Sagun, L., Biroli, G., and Wyart, M. (2018) · 2018
Later among the works it cites.
Exact expressions for double descent and implicit regularization via surrogate random design
Dereziński, M., Liang, F., and Mahoney, M. W. (2019) · 2019
Closest in time.
Jamming transition as a paradigm to understand the loss landscape of deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Belkin, M., Hsu, D., Ma, S., and Mandal, S. (2018) · 2018
Cited alongside, same era.
An analytic theory of generalization dynamics and transfer learning in deep linear networks
Lampinen, A. K. and Ganguli, S. (2018) · 2018
Cited alongside, same era.
Geiger, M., Spigler, S., d’Ascoli, S., Sagun, L., Baity-Jesi, M., Biroli, G., and Wyart, M. (2019) · 2019
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J. (2019) · 2019
Closest in time.
On the number of variables to use in principal component regression
Xu, J. and Hsu, D. J. (2019) · 2019
Closest in time.