Fetching the paper…
Reading the bibliography…
In many modern applications of deep learning the neural network has many more parameters than the data points used for its training.
Two models of double descent for weak features
M. Belkin, D. Hsu, and J. Xu · 1903
Earlier work this paper cites.
Linearized two-layers neural networks in high dimension
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 1904
Earlier work this paper cites.
When do neural networks outperform kernel methods?
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2006
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices , page 210–268
R. Vershynin · 2012
Earlier work this paper cites.
Random design analysis of ridge regression
D. Hsu, S. Kakade, T. Zhang, S. Mannor, N. Srebro, and B. Williamson · 2014
Earlier work this paper cites.
High-dimensional asymptotics of prediction: Ridge regression and classification
E. Dobriban and S. Wager · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Earlier work this paper cites.
On the interval of fluctuation of the singular values of random matrices
O. Guédon, A. Litvak, A. Pajor, and N. Tomczak-Jaegermann · 2017
Earlier work this paper cites.
Sample Covariance Matrices of Heavy-Tailed Distributions
K. Tikhomirov · 2017
Earlier work this paper cites.
Just interpolate: Kernel "ridgeless" regression can generalize
T. Liang and A. Rakhlin · 2018
Earlier work this paper cites.
High-Dimensional Probability: An Introduction with Applications in Data Science
R. Vershynin · 2018
Earlier work this paper cites.
A new look at an old problem: A universal learning approach to linear regression
K. Bibas, Y. Fogel, and M. Feder · 2019
Earlier work this paper cites.
Exact expressions for double descent and implicit regularization via surrogate random design
M. Dereziński, F. Liang, and M. Mahoney · 2019
Earlier work this paper cites.
Surprises in high-dimensional ridgeless least squares interpolation
T. J. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani · 2019
Earlier work this paper cites.
On the risk of minimum-norm interpolants and restricted lower isometry of kernels
T. Liang, A. Rakhlin, and X. Zhai · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and double descent curve
S. Mei and A. Montanari · 2019
Cited alongside, same era.
Harmless interpolation of noisy data in regression
V. Muthukumar, K. Vodrahalli, and A. Sahai · 2019
Cited alongside, same era.
More data can hurt for linear regression: Sample-wise double descent
P. Nakkiran · 2019
Cited alongside, same era.
On the number of variables to use in principal component regression
J. Xu and D. J. Hsu · 2019
Cited alongside, same era.
On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression
D. Wu and J. Xu · 2020
Closest in time.
Deep learning: a statistical viewpoint
P. L. Bartlett, A. Montanari, and A. Rakhlin · 2021
Closest in time.
Minimum complexity interpolation in random features models, 2021
M. Celentano, T. Misiakiewicz, and A. Montanari · 2021
Closest in time.
On the robustness of the minimum ℓ 2 \ell_{2} interpolator
G. Chinot and M. Lerasle · 2021
Closest in time.
The three stages of learning dynamics in high-dimensional kernel methods, 2021
N. Ghosh, S. Mei, and B. Yu · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benign overfitting in linear regression
P. L. Bartlett, P. M. Long, G. Lugosi, and A. Tsigler · 2020
Cited alongside, same era.
Precise expressions for random projections: Low-rank approximation and randomized newton
M. Dereziński, F. Liang, Z. Liao, and M. W. Mahoney · 2020
Cited alongside, same era.
Surprises in high-dimensional ridgeless least squares interpolation
T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani · 2020
Cited alongside, same era.
Optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization
D. Kobak, J. Lomond, and B. Sanchez · 2020
Cited alongside, same era.
On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels
T. Liang, A. Rakhlin, and X. Zhai · 2020
Cited alongside, same era.
A. Montanari and Y. Zhong · 2020
Cited alongside, same era.
In defense of uniform convergence: Generalization via derandomization with an application to interpolating predictors
J. Negrea, G. K. Dziugaite, and D. Roy · 2020
Cited alongside, same era.
Uniform convergence of interpolators: Gaussian width, norm bounds and benign overfitting
F. Koehler, L. Zhou, D. J. Sutherland, and N. Srebro · 2021
Closest in time.
Harmless interpolation in regression and classification with structured features, 2021
A. D. McRae, S. Karnik, M. A. Davenport, and V. Muthukumar · 2021
Closest in time.
Learning with convolution and pooling operations in kernel methods, 2021
T. Misiakiewicz and S. Mei · 2021
Closest in time.
Classification vs regression in overparameterized regimes: Does the loss function matter?
V. Muthukumar, A. Narang, V. Subramanian, M. Belkin, D. Hsu, and A. Sahai · 2021
Closest in time.
Classification and adversarial examples in an overparameterized linear model: A signal processing perspective, 2021
A. Narang, V. Muthukumar, and A. Sahai · 2021
Closest in time.
On uniform convergence and low-norm interpolation learning
L. Zhou, D. J. Sutherland, and N. Srebro · 2021
Closest in time.
Interpolating predictors in high-dimensional factor regression
F. Bunea, S. Strimas-Mackey, and M. H. Wegkamp · 2022
Closest in time.
The implicit bias of benign overfitting, 2022
O. Shamir · 2022
Closest in time.