Fetching the paper…
Reading the bibliography…
Deep neural networks generalize well despite being exceedingly overparameterized and being trained without explicit regularization.
Some inequalities for gaussian processes and applications
Y. Gordon · 1985
Earlier work this paper cites.
Linear and nonlinear extension of the pseudo-inverse solution for learning boolean functions
F. Vallet, J.-G. Cailton, and P. Refregier · 1989
Earlier work this paper cites.
On the ability of the optimal perceptron to generalise
M. Opper, W. Kinzel, J. Kleinz, and R. Nehl · 1990
Earlier work this paper cites.
Classifiers in almost empty spaces
R. P. Duin · 2000
Earlier work this paper cites.
Margin maximizing loss functions
S. Rosset, J. Zhu, and T. Hastie · 2003
Earlier work this paper cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?
V. Muthukumar, A. Narang, V. Subramanian, M. Belkin, D. Hsu, and A. Sahai · 2005
Earlier work this paper cites.
Introduction to algorithms
T. H. Cormen, C. E. Leiserson, R. L. Rivest, and C. Stein · 2009
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Earlier work this paper cites.
Impossibility of successful classification when useful features are rare and weak
J. Jin · 2009
Earlier work this paper cites.
Matrix analysis
R. A. Horn and C. R. Johnson · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
The squared-error of generalized lasso: A precise analysis
S. Oymak, C. Thrampoulidis, and B. Hassibi · 2013
Earlier work this paper cites.
Hanson-wright inequality and sub-gaussian concentration
M. Rudelson, R. Vershynin, et al · 2013
Earlier work this paper cites.
A framework to characterize performance of lasso algorithms
M. Stojnic · 2013
Earlier work this paper cites.
On the number of linear regions of deep neural networks
G. F. Montufar, R. Pascanu, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Earlier work this paper cites.
Regularized linear regression: A precise analysis of the estimation error
C. Thrampoulidis, S. Oymak, and B. Hassibi · 2015
Earlier work this paper cites.
Deep learning , volume 1
I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Cited alongside, same era.
Why and when can deep-but not shallow-networks avoid the curse of dimensionality: a review
T. Poggio, H. Mhaskar, L. Rosasco, B. Miranda, and Q. Liao · 2017
Cited alongside, same era.
D. Kobak, J. Lomond, and B. Sanchez · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro · 2018
Cited alongside, same era.
Precise error analysis of regularized m m -estimators in high dimensions
C. Thrampoulidis, E. Abbasi, and B. Hassibi · 2018
Cited alongside, same era.
Benign overfitting in linear regression
P. L. Bartlett, P. M. Long, G. Lugosi, and A. Tsigler · 2020
Closest in time.
X. Chang, Y. Li, S. Oymak, and C. Thrampoulidis · 2020
Closest in time.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
N. S. Chatterji and P. M. Long · 2020
Closest in time.
When does gradient descent with logistic loss find interpolating two-layer networks?
N. S. Chatterji, P. M. Long, and P. L. Bartlett · 2020
Closest in time.
On the proliferation of support vectors in high dimensions
D. Hsu, V. Muthukumar, and J. Xu · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High-dimensional probability: An introduction with applications in data science , volume 47
R. Vershynin · 2018
Cited alongside, same era.
Generalization of two-layer neural networks: An asymptotic viewpoint
J. Ba, M. Erdogdu, T. Suzuki, D. Wu, and T. Zhang · 2019
Cited alongside, same era.
Two models of double descent for weak features
M. Belkin, D. Hsu, and J. Xu · 2019
Cited alongside, same era.
A model of double descent for high-dimensional binary linear classification
Z. Deng, A. Kammoun, and C. Thrampoulidis · 2019
Cited alongside, same era.
Surprises in high-dimensional ridgeless least squares interpolation
T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani · 2019
Cited alongside, same era.
The implicit bias of gradient descent on nonseparable data
Z. Ji and M. Telgarsky · 2019
Cited alongside, same era.
On the risk of minimum-norm interpolants and restricted lower isometry of kernels
T. Liang, A. Rakhlin, and X. Zhai · 2019
Cited alongside, same era.
Closest in time.
On the precise error analysis of support vector machines
A. Kammoun and M.-S. Alouini · 2020
Closest in time.
Analytic study of double descent in binary classification: The impact of loss
G. Kini and C. Thrampoulidis · 2020
Closest in time.
Z. Liao, R. Couillet, and M. W. Mahoney · 2020
Closest in time.
A brief prehistory of double descent
M. Loog, T. Viering, A. Mey, J. H. Krijthe, and D. M. Tax · 2020
Closest in time.
The role of regularization in classification of high-dimensional noisy gaussian mixture
F. Mignacco, F. Krzakala, Y. M. Lu, and L. Zdeborová · 2020
Closest in time.
The performance analysis of generalized margin maximizer (gmm) on separable data
F. Salehi, E. Abbasi, and B. Hassibi · 2020
Closest in time.
Benign overfitting in ridge regression
A. Tsigler and P. L. Bartlett · 2020
Closest in time.
Rethinking bias-variance trade-off for generalization of neural networks
Z. Yang, Y. Yu, C. You, J. Steinhardt, and Y. Ma · 2020
Closest in time.
Support vector machines and linear regression coincide with very high-dimensional features
N. Ardeshir, C. Sanford, and D. Hsu · 2021
Closest in time.
Risk bounds for over-parameterized maximum margin classification on sub-gaussian mixtures
Y. Cao, Q. Gu, and M. Belkin · 2021
Closest in time.
Last iterate convergence of sgd for least-squares in the interpolation regime
A. Varre, L. Pillaud-Vivien, and N. Flammarion · 2021
Closest in time.
Benign overfitting in binary classification of gaussian mixtures
K. Wang and C. Thrampoulidis · 2021
Closest in time.
Benign overfitting of constant-stepsize sgd for linear regression
D. Zou, J. Wu, V. Braverman, Q. Gu, and S. M. Kakade · 2021
Closest in time.