Fetching the paper…
Reading the bibliography…
Recently, there have been significant interests in studying the so-called "double-descent" of the generalization error of linear regression models under the overparameterized and overfitting regime, with the hope that such analysis may provide the first step towards understanding why overparameterized deep neural networks (DNN) still generalize well.
On the stability of inverse problems
A. N. Tikhonov · 1943
Earlier work this paper cites.
Inadmissibility of the usual estimator for the mean of a multivariate normal distribution
C. Stein · 1956
Earlier work this paper cites.
A problem in geometric probability
J. G. Wendel · 1962
Earlier work this paper cites.
Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables. National Bureau of Standards Applied Mathematics Series 55. Tenth Printing
M. Abramowitz and I. A. Stegun · 1972
Earlier work this paper cites.
Occam’s razor
A. Blumer, A. Ehrenfeucht, D. Haussler, and M. K. Warmuth · 1987
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Second order properties of error surfaces: Learning time and generalization
Y. LeCun, I. Kanter, and S. A. Solla · 1991
Earlier work this paper cites.
Estimation with quadratic loss
W. James and C. Stein · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
A. R. Barron · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
A. R. Barron · 1994
Earlier work this paper cites.
Examples of basis pursuit
S. Chen and D. L. Donoho · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the Lasso
R. Tibshirani · 1996
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
P. L. Bartlett · 1998
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
B. Laurent and P. Massart · 2000
Cited alongside, same era.
Atomic decomposition by basis pursuit
S. S. Chen, D. L. Donoho, and M. A. Saunders · 2001
Cited alongside, same era.
Uncertainty principles and ideal atomic decomposition
D. L. Donoho and X. Huo · 2001
Cited alongside, same era.
Stable recovery of sparse overcomplete representations in the presence of noise
D. L. Donoho, M. Elad, and V. Temlyakov · 2004
Cited alongside, same era.
Stable recovery of sparse overcomplete representations in the presence of noise
D. L. Donoho, M. Elad, and V. N. Temlyakov · 2005
Cited alongside, same era.
Pattern recognition and machine learning
C. M. Bishop · 2006
Cited alongside, same era.
Benefits of depth in neural networks
M. Telgarsky · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Later among the works it cites.
To understand deep learning we need to understand kernel learning
M. Belkin, S. Ma, and S. Mandal · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P.-M. Nguyen · 2018
Later among the works it cites.
Two models of double descent for weak features
M. Belkin, D. Hsu, and J. Xu · 2019
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Zhao and B. Yu · 2006
Cited alongside, same era.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Cited alongside, same era.
Simultaneous analysis of Lasso and Dantzig selector
P. J. Bickel, Y. Ritov, and A. B. Tsybakov · 2009
Cited alongside, same era.
The elements of statistical learning: data mining, inference, and prediction
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Cited alongside, same era.
Lasso-type recovery of sparse representations for high-dimensional data
N. Meinshausen and B. Yu · 2009
Cited alongside, same era.
Chernoff bounds, and some applications
M. Goemans · 2015
Cited alongside, same era.
T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
S. Mei and A. Montanari · 2019
Later among the works it cites.
P. P. Mitra · 2019
Later among the works it cites.
Harmless interpolation of noisy data in regression
V. Muthukumar, K. Vodrahalli, and A. Sahai · 2019
Later among the works it cites.
High-dimensional dynamics of generalization error in neural networks
M. S. Advani, A. M. Saxe, and H. Sompolinsky · 2020
Closest in time.
Benign overfitting in linear regression
P. L. Bartlett, P. M. Long, G. Lugosi, and A. Tsigler · 2020
Closest in time.
T. Liang and P. Sur · 2020
Closest in time.