Fetching the paper…
Reading the bibliography…
A conventional wisdom in statistical learning is that large models require strong regularization to prevent overfitting.
Two models of double descent for weak features
M. Belkin, D. Hsu, and J. Xu · 1903
Earlier work this paper cites.
On the solution of ill-posed problems and the method of regularization
A. N. Tikhonov · 1963
Earlier work this paper cites.
Ridge regression: Biased estimation for nonorthogonal problems
A. E. Hoerl and R. W. Kennard · 1970
Earlier work this paper cites.
Training with noise is equivalent to Tikhonov regularization
C. M. Bishop · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
The Nature of Statistical Learning Theory
V. Vapnik · 1996
Earlier work this paper cites.
Atomic decomposition by basis pursuit
S. S. Chen, D. L. Donoho, and M. A. Saunders · 2001
Earlier work this paper cites.
Optimal regularization can mitigate double descent
P. Nakkiran, P. Venkat, S. Kakade, and T. Ma · 2003
Earlier work this paper cites.
Regularization and variable selection via the elastic net
H. Zou and T. Hastie · 2005
Earlier work this paper cites.
Simultaneous clustering of gene expression data with clinical chemistry and pathological evaluations reveals phenotypic prototypes
P. R. Bushel, R. D. Wolfinger, and G. Gibson · 2007
Earlier work this paper cites.
The Dantzig selector: Statistical estimation when p p is much larger than n n
E. Candes and T. Tao · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Earlier work this paper cites.
The Elements of Statistical Learning
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Cited alongside, same era.
Regularization paths for generalized linear models via coordinate descent
J. Friedman, T. Hastie, and R. Tibshirani · 2010
Cited alongside, same era.
An Introduction to Statistical Learning , volume 112
G. James, D. Witten, T. Hastie, and R. Tibshirani · 2013
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Statistical Learning with Sparsity: the Lasso and Generalizations
T. Hastie, R. Tibshirani, and M. Wainwright · 2015
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
High-dimensional asymptotics of prediction: Ridge regression and classification
E. Dobriban and S. Wager · 2018
Closest in time.
Sparse reduced-rank regression for exploratory visualization of multimodal data sets
D. Kobak, Y. Bernaerts, M. A. Weis, F. Scala, A. Tolias, and P. Berens · 2018
Closest in time.
Just interpolate: Kernel “ridgeless” regression can generalize
T. Liang and A. Rakhlin · 2018
Closest in time.
Benign overfitting in linear regression
P. L. Bartlett, P. M. Long, G. Lugosi, and A. Tsigler · 2019
Closest in time.
A new look at an old problem: A universal learning approach to linear regression
K. Bibas, Y. Fogel, and M. Feder · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. S. Advani and A. M. Saxe · 2017
Cited alongside, same era.
Implicit Regularization in Deep Learning
B. Neyshabur · 2017
Cited alongside, same era.
Theory of deep learning III: explaining the non-overfitting puzzle
T. Poggio, K. Kawaguchi, Q. Liao, B. Miranda, L. Rosasco, X. Boix, J. Hidary, and H. Mhaskar · 2017
Cited alongside, same era.
mixOmics: An R package for ‘omics feature selection and multiple data integration
F. Rohart, B. Gautier, A. Singh, and K.-A. Le Cao · 2017
Cited alongside, same era.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, and N. Srebro · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Cited alongside, same era.
Exact expressions for double descent and implicit regularization via surrogate random design
M. Dereziński, F. Liang, and M. W. Mahoney · 2019
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation
T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani · 2019
Closest in time.
The generalization error of random features regression: Precise asymptotics and double descent curve
S. Mei and A. Montanari · 2019
Closest in time.
Harmless interpolation of noisy data in regression
V. Muthukumar, K. Vodrahalli, and A. Sahai · 2019
Closest in time.
More data can hurt for linear regression: Sample-wise double descent
P. Nakkiran · 2019
Closest in time.
A jamming transition from under-to over-parametrization affects generalization in deep learning
S. Spigler, M. Geiger, S. d’Ascoli, L. Sagun, G. Biroli, and M. Wyart · 2019
Closest in time.
Benign overfitting in the large deviation regime
G. Chinot and M. Lerasle · 2020
Closest in time.