Fetching the paper…
Reading the bibliography…
Overparameterized neural networks (NNs) are observed to generalize well even when trained to perfectly fit noisy data.
Spacings
Ronald Pyke · 1965
Earlier work this paper cites.
Nearest neighbor pattern classification
Thomas Cover and Peter Hart · 1967
Earlier work this paper cites.
On the lengths of the pieces of a stick broken at random
Lars Holst · 1980
Earlier work this paper cites.
Optimization and Nonsmooth Analysis
F. H. Clarke · 1990
Earlier work this paper cites.
An elementary introduction to modern convex geometry
Keith Ball · 1997
Earlier work this paper cites.
Approximate kkt points and a proximity measure for termination
Joydeep Dutta, Kalyanmoy Deb, Rupesh Tulshyan, and Ramnik Arora · 2013
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin · 2018
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan · 2019
Earlier work this paper cites.
Uniform convergence may be unable to explain generalization in deep learning
Vaishnavh Nagarajan and J Zico Kolter · 2019
Earlier work this paper cites.
Generalization in deep network classifiers trained with the square loss
Tomaso Poggio and Qianli Liao · 2019
Earlier work this paper cites.
Consistency of interpolation with laplace kernels is a high-dimensional phenomenon
Alexander Rakhlin and Xiyu Zhai · 2019
Earlier work this paper cites.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Earlier work this paper cites.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2020
Earlier work this paper cites.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Earlier work this paper cites.
Just interpolate: Kernel “ridgeless” regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2020
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Cited alongside, same era.
Distributional generalization: A new kind of generalization
Preetum Nakkiran and Yamini Bansal · 2020
Cited alongside, same era.
In defense of uniform convergence: Generalization via derandomization with an application to interpolating predictors
Jeffrey Negrea, Gintare Karolina Dziugaite, and Daniel Roy · 2020
Cited alongside, same era.
Theoretical insights into multiclass classification: A high-dimensional asymptotic view
Christos Thrampoulidis, Samet Oymak, and Mahdi Soltanolkotabi · 2020
Cited alongside, same era.
Uniform convergence, adversarial spheres and a simple remedy
Gregor Bachmann, Seyed-Mohsen Moosavi-Dezfooli, and Thomas Hofmann · 2021
Cited alongside, same era.
Understanding square loss in training overparametrized neural network classifiers
Tianyang Hu, Jun Wang, Wenjia Wang, and Zhenguo Li · 2022
Later among the works it cites.
Benign, tempered, or catastrophic: Toward a refined taxonomy of overfitting
Neil Mallinar, James Simon, Amirhesam Abedsoltan, Parthe Pandit, Misha Belkin, and Preetum Nakkiran · 2022
Later among the works it cites.
Harmless interpolation in regression and classification with structured features
Andrew D McRae, Santhosh Karnik, Mark Davenport, and Vidya K Muthukumar · 2022
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Later among the works it cites.
On the effective number of linear regions in shallow univariate relu networks: Convergence guarantees and implicit bias
Itay Safran, Gal Vardi, and Jason D Lee · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Failures of model-dependent generalization bounds for least-norm interpolation
Peter L Bartlett and Philip M Long · 2021
Cited alongside, same era.
Deep learning: a statistical viewpoint
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin · 2021
Cited alongside, same era.
Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
Mikhail Belkin · 2021
Cited alongside, same era.
Risk bounds for over-parameterized maximum margin classification on sub-gaussian mixtures
Yuan Cao, Quanquan Gu, and Mikhail Belkin · 2021
Cited alongside, same era.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
Niladri S Chatterji and Philip M Long · 2021
Cited alongside, same era.
Uniform convergence of interpolators: Gaussian width, norm bounds and benign overfitting
Frederic Koehler, Lijia Zhou, Danica J Sutherland, and Nathan Srebro · 2021
Cited alongside, same era.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu, and Anant Sahai · 2021
Cited alongside, same era.
Gal Vardi · 2022
Later among the works it cites.
On margin maximization in linear and relu networks
Gal Vardi, Ohad Shamir, and Nati Srebro · 2022
Later among the works it cites.
On the inconsistency of kernel ridgeless regression in fixed dimensions
Daniel Beaglehole, Mikhail Belkin, and Parthe Pandit · 2023
Closest in time.
Penalising the biases in norm regularisation enforces sparsity
Etienne Boursier and Nicolas Flammarion · 2023
Closest in time.
Deep linear networks can benignly overfit when shallow ones do
Niladri S Chatterji and Philip M Long · 2023
Closest in time.
Spencer Frei, Gal Vardi, Peter L Bartlett, and Nathan Srebro · 2023
Closest in time.
Noisy interpolation learning with shallow univariate relu networks
Nirmit Joshi, Gal Vardi, and Nathan Srebro · 2023
Closest in time.
Benign overfitting for two-layer relu networks
Yiwen Kou, Zixiang Chen, Yuanzhou Chen, and Quanquan Gu · 2023
Closest in time.
Interpolating classifiers make few mistakes
Tengyuan Liang and Benjamin Recht · 2023
Closest in time.
Interpolation learning with minimum description length
Naren Sarayu Manoj and Nathan Srebro · 2023
Closest in time.
The implicit bias of benign overfitting
Ohad Shamir · 2023
Closest in time.
Benign overfitting of non-smooth neural networks beyond lazy training
Xingyu Xu and Yuantao Gu · 2023
Closest in time.
Lijia Zhou, Frederic Koehler, Danica J Sutherland, and Nathan Srebro · 2023
Closest in time.