Fetching the paper…
Reading the bibliography…
Benign overfitting, the phenomenon where interpolating models generalize well in the presence of noisy data, was first observed in neural network models trained with gradient descent.
“Structural risk minimization over data-dependent hierarchies”
John Shawe-Taylor, Peter Bartlett, Robert Williamson and Martin Anthony · 1940
Earlier work this paper cites.
“Toward efficient agnostic learning”
Michael Kearns, Robert Schapire and Linda Sellie · 1994
Earlier work this paper cites.
“The nature of statistical learning theory”
Vladimir Vapnik · 1999
Earlier work this paper cites.
“The Concentration of Measure Phenomenon”, Mathematical surveys and monographs
M. Ledoux · 2001
Earlier work this paper cites.
“Introduction to the non-asymptotic analysis of random matrices”
Roman Vershynin · 2010
Earlier work this paper cites.
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht and Oriol Vinyals · 2017
Earlier work this paper cites.
“Overfitting or perfect fitting? Risk bounds for classification and regression rules that interpolate”
Mikhail Belkin, Daniel Hsu and Partha Mitra · 2018
Earlier work this paper cites.
“SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data”
Alon Brutzkus, Amir Globerson, Eran Malach and Shai Shalev-Shwartz · 2018
Earlier work this paper cites.
“Neural Tangent Kernel: Convergence and Generalization in Neural Networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Earlier work this paper cites.
“The Implicit Bias of Gradient Descent on Separable Data”
Daniel Soudry, Elad Hoffer, Mor Nacson, Suriya Gunasekar and Nathan Srebro · 2018
Earlier work this paper cites.
“A Convergence Theory for Deep Learning via Over-Parameterization”
Zeyuan Allen-Zhu, Yuanzhi Li and Zhao Song · 2019
Earlier work this paper cites.
“On exact computation with an infinitely wide neural net”
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov and Ruosong Wang · 2019
Earlier work this paper cites.
“Reconciling modern machine-learning practice and the classical bias–variance trade-off”
Mikhail Belkin, Daniel Hsu, Siyuan Ma and Soumik Mandal · 2019
Earlier work this paper cites.
“Gradient Descent Provably Optimizes Over-parameterized Neural Networks”
Simon. Du, Xiyu Zhai, Barnabás Póczos and Aarti Singh · 2019
Earlier work this paper cites.
“Algorithm-Dependent Generalization Bounds for Overparameterized Deep Residual Networks”
Spencer Frei, Yuan Cao and Quanquan Gu · 2019
Earlier work this paper cites.
“Partial recovery bounds for clustering with the relaxed K K -means”
Christophe Giraud and Nicolas Verzelen · 2019
Earlier work this paper cites.
“Surprises in high-dimensional ridgeless least squares interpolation”
Trevor Hastie, Andrea Montanari, Saharon Rosset and Ryan Tibshirani · 2019
Earlier work this paper cites.
“The implicit bias of gradient descent on nonseparable data”
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
“The generalization error of random features regression: Precise asymptotics and the double descent curve”
Song Mei and Andrea Montanari · 2019
Cited alongside, same era.
Andrea Montanari, Feng Ruan, Youngtak Sohn and Jun Yan · 2019
Cited alongside, same era.
“Theoretical Insights Into the Optimization Landscape of Over-Parameterized Shallow Neural Networks”
Mahdi Soltanolkotabi, Adel Javanmard and Jason. Lee · 2019
Cited alongside, same era.
“High-Dimensional Statistics: A Non-Asymptotic Viewpoint”, Cambridge Series in Statistical and Probabilistic Mathematics
M.J. Wainwright · 2019
Cited alongside, same era.
“Gradient descent optimizes over-parameterized deep ReLU networks”
Difan Zou, Yuan Cao, Dongruo Zhou and Quanquan Gu · 2019
“Finite-sample analysis of interpolating linear classifiers in the overparameterized regime”
Niladri. Chatterji and Philip. Long · 2021
Later among the works it cites.
“Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label Noise”
Spencer Frei, Yuan Cao and Quanquan Gu · 2021
Later among the works it cites.
“Proxy Convexity: A Unified Framework for the Analysis of Neural Networks Trained by Gradient Descent”
Spencer Frei and Quanquan Gu · 2021
Later among the works it cites.
“On the proliferation of support vectors in high dimensions”
Daniel Hsu, Vidya Muthukumar and Ji Xu · 2021
Later among the works it cites.
“Early-stopped neural networks are consistent”
Ziwei Ji, Justin. Li and Matus Telgarsky · 2021
Later among the works it cites.
“Uniform convergence of interpolators: Gaussian width, norm bounds, and benign overfitting”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Benign Overfitting in Linear Regression”
Peter. Bartlett, Philip. Long, Gábor Lugosi and Alexander Tsigler · 2020
Cited alongside, same era.
“Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks”
Yuan Cao and Quanquan Gu · 2020
Cited alongside, same era.
“On the robustness of minimum-norm interpolators”
Geoffrey Chinot, Matthias Löffler and Sara van Geer · 2020
Cited alongside, same era.
“Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks”
Ziwei Ji and Matus Telgarsky · 2020
Cited alongside, same era.
“Just interpolate: Kernel “ridgeless” regression can generalize”
Tengyuan Liang and Alexander Rakhlin · 2020
Cited alongside, same era.
“On the Multiple Descent of Minimum-Norm Interpolants and Restricted Lower Isometry of Kernels”
Tengyuan Liang, Alexander Rakhlin and Xiyu Zhai · 2020
Cited alongside, same era.
Tengyuan Liang and Pragya Sur · 2020
Cited alongside, same era.
Frederic Koehler, Lijia Zhou, Danica Sutherland and Nathan Srebro · 2021
Later among the works it cites.
“Minimum ℓ 1 \ell_{1} -norm interpolators: Precise asymptotics and multiple descent”
Yue Li and Yuting Wei · 2021
Later among the works it cites.
“Interpolating classifiers make few mistakes”
Tengyuan Liang and Benjamin Recht · 2021
Later among the works it cites.
Stanislav Minsker, Mohamed Ndaoud and Yiqiu Shen · 2021
Later among the works it cites.
Andrea Montanari and Yiqiao Zhong · 2021
Later among the works it cites.
“Classification vs regression in overparameterized regimes: Does the loss function matter?”
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu and Anant Sahai · 2021
Later among the works it cites.
“Tight bounds for minimum l1-norm interpolation of noisy data”
Guillaume Wang, Konstantin Donhauser and Fanny Yang · 2021
Later among the works it cites.
“Benign Overfitting in Multiclass Classification: All Roads Lead to Interpolation”
Ke Wang, Vidya Muthukumar and Christos Thrampoulidis · 2021
Later among the works it cites.
Ke Wang and Christos Thrampoulidis · 2021
Later among the works it cites.
“Is Importance Weighting Incompatible with Interpolating Classifiers?”
Ke Wang, Niladri Chatterji, Saminul Haque and Tatsunori Hashimoto · 2021
Later among the works it cites.
“Feature Learning in Infinite-Width Neural Networks”
Greg Yang and Edward. Hu · 2021
Later among the works it cites.
“Benign Overfitting in Two-layer Convolutional Neural Networks”
Yuan Cao, Zixiang Chen, Mikhail Belkin and Quanquan Gu · 2022
Closest in time.