Fetching the paper…
Reading the bibliography…
Recent theoretical work has guaranteed that overparameterized networks trained by gradient descent achieve arbitrarily low training error, and sometimes even low test error.
Can sgd learn recurrent neural networks with provable generalization?
Zeyuan Allen-Zhu and Yuanzhi Li · 1902
Earlier work this paper cites.
Generalization error bounds of gradient descent for learning over-parameterized deep relu networks
Yuan Cao and Quanquan Gu · 1902
Earlier work this paper cites.
What can resnet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 1905
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 1905
Earlier work this paper cites.
On convergence proofs on perceptrons
Albert B.J. Novikoff · 1962
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Techniques of Variational Analysis, volume 20 of
Jonathan M. Borwein and Qiji J. Zhu · 2005
Earlier work this paper cites.
Smoothness, low noise and fast rates
Nathan Srebro, Karthik Sridharan, and Ambuj Tewari · 2010
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Cited alongside, same era.
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
UC Berkeley Statistics 210B, Lecture Notes: Basic tail and concentration bounds, Jan 2015
Martin J. Wainwright · 2015
Cited alongside, same era.
Stanford CS229T/STAT231: Statistical Learning Theory, Apr 2016
Percy Liang · 2016
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Later among the works it cites.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Closest in time.
A Note on Lazy Training in Supervised Differentiable Programming
Lenaic Chizat and Francis Bach · 2019
Closest in time.
Atsushi Nitanda and Taiji Suzuki · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ziwei Ji and Matus Telgarsky · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang
Cited in the paper.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song
Cited in the paper.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song
Cited in the paper.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai
Cited in the paper.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Closest in time.
Quadratic suffices for over-parametrization via matrix chernoff bound
Zhao Song and Xin Yang · 2019
Closest in time.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Closest in time.