Fetching the paper…
Reading the bibliography…
This paper explores the generalization loss of linear regression in variably parameterized families of models, both under-parameterized and over-parameterized.
Two models of double descent for weak features
M. Belkin, D. Hsu, and J. Xu · 1903
Earlier work this paper cites.
Hidden integrality of sdp relaxations for sub-gaussian mixture models
Y. Fei and Y. Chen · 1965
Earlier work this paper cites.
Moments for the inverted wishart distribution
D. von Rosen · 1988
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
S. Geman, E. Bienenstock, and R. Doursat · 1992
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin · 2003
Earlier work this paper cites.
Model selection for regularized least-squares algorithm in learning theory
E. De Vito, A. Caponnetto, and L. Rosasco · 2005
Earlier work this paper cites.
Particular formulae for the moore–penrose inverse of a columnwise partitioned matrix
J. K. Baksalary and O. M. Baksalary · 2007
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
A. Caponnetto and E. De Vito · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
M. S. Advani and A. M. Saxe · 2017
Earlier work this paper cites.
Generalization properties of learning with random features
A. Rudi and L. Rosasco · 2017
Earlier work this paper cites.
Explaining the success of adaboost and random forests as interpolating classifiers
A. J. Wyner, M. Olson, J. Bleich, and D. Mease · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
M. Caron, P. Bojanowski, A. Joulin, and M. Douze · 2018
Earlier work this paper cites.
A modern take on the bias-variance tradeoff in neural networks
B. Neal, S. Mittal, A. Baratin, V. Tantia, M. Scicluna, S. Lacoste-Julien, and I. Mitliagkas · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Y. Cao and Q. Gu · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
S. Du, J. Lee, H. Li, L. Wang, and X. Zhai · 2019
Cited alongside, same era.
Achieving the bayes error rate in stochastic block model by sdp, robustly
Y. Fei and Y. Chen · 2019
Cited alongside, same era.
Jamming transition as a paradigm to understand the loss landscape of deep neural networks
A finite sample analysis of the double descent phenomenon for ridge function estimation
E. Caron and S. Chretien · 2020
Closest in time.
More data can expand the generalization gap between adversarially robust and standard models
L. Chen, Y. Min, M. Zhang, and A. Karbasi · 2020
Closest in time.
Subspace fitting meets regression: The effects of supervision and orthonormality constraints on double descent of generalization errors
Y. Dar, P. Mayer, L. Luzi, and R. G. Baraniuk · 2020
Closest in time.
Triple descent and the two kinds of overfitting: Where & why do they appear?
S. d’Ascoli, L. Sagun, and G. Biroli · 2020
Closest in time.
Achieving the bayes error rate in synchronization and block models by sdp, robustly
Y. Fei and Y. Chen · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Geiger, S. Spigler, S. d’Ascoli, L. Sagun, M. Baity-Jesi, G. Biroli, and M. Wyart · 2019
Cited alongside, same era.
Linearized two-layers neural networks in high dimension
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Cited alongside, same era.
Surprises in high-dimensional ridgeless least squares interpolation
T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, et al · 2019
Cited alongside, same era.
Just interpolate: Kernel “ridgeles” regression can generalize
T. Liang and A. Rakhlin · 2019
Cited alongside, same era.
Minimizers of the empirical risk and risk monotonicity
M. Loog, T. Viering, and A. Mey · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and double descent curve
S. Mei and A. Montanari · 2019
Cited alongside, same era.
Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret
Y. Fei, Z. Yang, Y. Chen, Z. Wang, and Q. Xie · 2020
Closest in time.
Scaling description of generalization with number of parameters in deep learning
M. Geiger, A. Jacot, S. Spigler, F. Gabriel, L. Sagun, S. d’Ascoli, G. Biroli, C. Hongler, and M. Wyart · 2020
Closest in time.
Precise tradeoffs in adversarial training for linear regression
A. Javanmard, M. Soltanolkotabi, and H. Hassani · 2020
Closest in time.
On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels
T. Liang, A. Rakhlin, and X. Zhai · 2020
Closest in time.
Y. Min, L. Chen, and A. Karbasi · 2020
Closest in time.
Optimal regularization can mitigate double descent
P. Nakkiran, P. Venkat, S. Kakade, and T. Ma · 2020
Closest in time.
Asymptotics of ridge (less) regression under general source condition
D. Richards, J. Mourtada, and L. Rosasco · 2020
Closest in time.
Benign overfitting in ridge regression
A. Tsigler and P. L. Bartlett · 2020
Closest in time.
Gradient descent optimizes over-parameterized deep relu networks
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2020
Closest in time.
Deep neural tangent kernel and laplace kernel have the same rkhs
L. Chen and S. Xu · 2021
Closest in time.
Minimum ℓ 1 \ell_{1} -norm interpolators: Precise asymptotics and multiple descent
Y. Li and Y. Wei · 2021
Closest in time.
Kernel regression in high dimensions: Refined analysis beyond double descent
F. Liu, Z. Liao, and J. Suykens · 2021
Closest in time.
Convergence and alignment of gradient descent with random back propagation weights
G. Song, R. Xu, and J. Lafferty · 2021
Closest in time.