Fetching the paper…
Reading the bibliography…
Modern machine learning models often employ a huge number of parameters and are typically optimized to have zero training loss; yet surprisingly, they possess near-optimal prediction performance, contradicting classical learning theory.
Two models of double descent for weak features
M. Belkin, D. Hsu, and J. Xu · 1903
Earlier work this paper cites.
Training with noise is equivalent to tikhonov regularization
C. M. Bishop · 1995
Earlier work this paper cites.
The elements of statistical learning , volume 1
J. Friedman, T. Hastie, and R. Tibshirani · 2001
Earlier work this paper cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?
V. Muthukumar, A. Narang, V. Subramanian, M. Belkin, D. Hsu, and A. Sahai · 2005
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
A. Caponnetto and E. De Vito · 2007
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
An introduction to matrix concentration inequalities
J. A. Tropp · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Earlier work this paper cites.
Random Fourier features for kernel ridge regression: Approximation bounds and statistical guarantees
H. Avron, M. Kapralov, C. Musco, C. Musco, A. Velingker, and A. Zandieh · 2017
Earlier work this paper cites.
To understand deep learning we need to understand kernel learning
M. Belkin, S. Ma, and S. Mandal · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science , volume 47
R. Vershynin · 2018
Earlier work this paper cites.
A model of double descent for high-dimensional binary linear classification
Z. Deng, A. Kammoun, and C. Thrampoulidis · 2019
Earlier work this paper cites.
Exact expressions for double descent and implicit regularization via surrogate random design
M. Dereziński, F. Liang, and M. W. Mahoney · 2019
Cited alongside, same era.
Surprises in high-dimensional ridgeless least squares interpolation
T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani · 2019
Cited alongside, same era.
Risk of the least squares minimum norm estimator under the spike covariance model
Y. Mahdaviyeh and Z. Naulet · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and double descent curve
S. Mei and A. Montanari · 2019
Cited alongside, same era.
Ivanov-regularised least-squares estimators over large rkhss and their interpolation spaces
Triple descent and the two kinds of overfitting: Where & why do they appear?
S. d’Ascoli, L. Sagun, and G. Biroli · 2020
Later among the works it cites.
Implicit regularization of random feature models
A. Jacot, B. Simsek, F. Spadaro, C. Hongler, and F. Gabriel · 2020
Later among the works it cites.
Benign overfitting and noisy features
Z. Li, W. Su, and D. Sejdinovic · 2020
Later among the works it cites.
A random matrix analysis of random fourier features: beyond the gaussian kernel, a precise phase transition, and the corresponding double descent
Z. Liao, R. Couillet, and M. Mahoney · 2020
Later among the works it cites.
Optimal regularization can mitigate double descent
P. Nakkiran, P. Venkat, S. Kakade, and T. Ma · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Page and S. Grünewälder · 2019
Cited alongside, same era.
Consistency of interpolation with laplace kernels is a high-dimensional phenomenon
A. Rakhlin and X. Zhai · 2019
Cited alongside, same era.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
B. Adlam and J. Pennington · 2020
Cited alongside, same era.
Benign overfitting in linear regression
P. L. Bartlett, P. M. Long, G. Lugosi, and A. Tsigler · 2020
Cited alongside, same era.
Interpolation under latent factor regression models
F. Bunea, S. Strimas-Mackey, and M. Wegkamp · 2020
Cited alongside, same era.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
N. S. Chatterji and P. M. Long · 2020
Cited alongside, same era.
Multiple descent: Design your own generalization curve
L. Chen, Y. Min, M. Belkin, and A. Karbasi · 2020
Cited alongside, same era.
Benign overfitting in the large deviation regime
G. Chinot and M. Lerasle · 2020
Cited alongside, same era.
Later among the works it cites.
Benign overfitting in binary classification of gaussian mixtures
K. Wang and C. Thrampoulidis · 2020
Later among the works it cites.
On the optimal weighted l 2 l_{2} regularization in overparameterized linear regression
D. Wu and J. Xu · 2020
Later among the works it cites.
Risk bounds for over-parameterized maximum margin classification on sub-gaussian mixtures
Y. Cao, Q. Gu, and M. Belkin · 2021
Closest in time.
Linearized two-layers neural networks in high dimension
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2021
Closest in time.
Kernel regression in high dimensions: Refined analysis beyond double descent
F. Liu, Z. Liao, and J. Suykens · 2021
Closest in time.
Asymptotics of ridge (less) regression under general source condition
D. Richards, J. Mourtada, and L. Rosasco · 2021
Closest in time.
Why over-parameterization of deep neural networks does not overfit?
Z.-H. Zhou · 2021
Closest in time.