Fetching the paper…
Reading the bibliography…
We revisit on-average algorithmic stability of GD for training overparameterised shallow neural networks and prove new generalisation and excess risk bounds without the NTK or PL assumptions.
Neural network learning: Theoretical foundations
M. Anthony and P. L. Bartlett · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
B. Schölkopf and A. J. Smola · 2002
Earlier work this paper cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2003
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2003
Earlier work this paper cites.
On early stopping in gradient descent learning
Y. Yao, L. Rosasco, and A. Caponnetto · 2007
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
L. Sagun, L. Bottou, and Y. LeCun · 2016
Earlier work this paper cites.
Fast rates for empirical risk minimization of strict saddle problems
A. Gonen and S. Shalev-Shwartz · 2017
Earlier work this paper cites.
A second-order look at stability and generalization
A. Maurer · 2017
Earlier work this paper cites.
Stability and generalization of learning algorithms that converge to global optima
Z. Charles and D. Papailiopoulos · 2018
Earlier work this paper cites.
Stability and convergence trade-off of iterative optimization algorithms
Y. Chen, C. Jin, and B. Yu · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2018
Cited alongside, same era.
Size-independent sample complexity of neural networks
N. Golowich, A. Rakhlin, and O. Shamir · 2018
Cited alongside, same era.
Neural tangent kernel: convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Data-dependent stability of stochastic gradient descent
I. Kuzborskij and C. H. Lampert · 2018
Cited alongside, same era.
The role of over-parametrization in generalization of neural networks
Sharper bounds for uniformly stable algorithms
O. Bousquet, Y. Klochkov, and N. Zhivotovskiy · 2020
Later among the works it cites.
Fine-grained analysis of stability and generalization for stochastic gradient descent
Y. Lei and Y. Ying · 2020
Later among the works it cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
M. Li, M. Soltanolkotabi, and S. Oymak · 2020
Later among the works it cites.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
S. Oymak and M. Soltanolkotabi · 2020
Later among the works it cites.
For interpolating kernel machines, minimizing the norm of the ERM solution minimizes stability
A. Rangamani, L. Rosasco, and T. Poggio · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Neyshabur, Z. Li, S. Bhojanapalli, Y. LeCun, and N. Srebro · 2018
Cited alongside, same era.
An exponential efron-stein inequality for l _ q l\_q stable learning rules
K. Abou-Moustafa and Cs. Szepesvári · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
S. Arora, S. Du, W. Hu, Z. Li, and R. Wang · 2019
Cited alongside, same era.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Y. Bai and J. D. Lee · 2019
Cited alongside, same era.
High probability generalization bounds for uniformly stable algorithms with nearly optimal rate
V. Feldman and J. Vondrak · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Cited alongside, same era.
M. Seleznova and G. Kutyniok · 2020
Later among the works it cites.
Deep learning: a statistical viewpoint
P. L. Bartlett, A. Montanari, and A. Rakhlin · 2021
Closest in time.
Regularization matters: A nonparametric perspective on overparametrized neural network
T. Hu, W. Wang, C. Lin, and G. Cheng · 2021
Closest in time.
Early-stopped neural networks are consistent
Z. Ji, J. D. Li, and M. Telgarsky · 2021
Closest in time.
Nonparametric regression with shallow overparameterized neural networks trained by GD with early stopping
I. Kuzborskij and Cs. Szepesvári · 2021
Closest in time.
Sharper generalization bounds for learning with gradient-dominated objective functions
Y. Lei and Y. Ying · 2021
Closest in time.
Learning with gradient descent and weakly convex losses
D. Richards and M. Rabbat · 2021
Closest in time.
Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods
T. Suzuki and S. Akiyama · 2021
Closest in time.