Fetching the paper…
Reading the bibliography…
Recently, several studies have proven the global convergence and generalization abilities of the gradient descent method for two-layer ReLU networks.
A generalization theory of gradient descent for learning over-parameterized deep relu networks
Yuan Cao and Quanquan Gu · 1902
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 1905
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
Boosting algorithms as gradient descent
Llew Mason, Jonathan Baxter, Peter L Bartlett, and Marcus R Frean · 1999
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Jerome H Friedman · 2001
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
Vladimir Koltchinskii and Dmitry Panchenko · 2002
Earlier work this paper cites.
Online learning with kernels
Jyrki Kivinen, Alexander J Smola, and Robert C Williamson · 2004
Earlier work this paper cites.
Online learning algorithms
Steve Smale and Yuan Yao · 2006
Earlier work this paper cites.
Online regularized classification algorithms
Yiming Ying and D-X Zhou · 2006
Earlier work this paper cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2012
Earlier work this paper cites.
Early stopping and non-parametric regression: an optimal data-dependent stopping rule
Garvesh Raskutti, Martin J Wainwright, and Bin Yu · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
Stochastic particle gradient descent for infinite ensembles
Atsushi Nitanda and Taiji Suzuki · 2017
Cited alongside, same era.
Learning relus via gradient descent
Mahdi Soltanolkotabi · 2017
Cited alongside, same era.
Neural network with unbounded activation functions is universal approximator
Sho Sonoda and Noboru Murata · 2017
Cited alongside, same era.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2018
Later among the works it cites.
Learning one-hidden-layer relu networks via gradient descent
Xiao Zhang, Yaodong Yu, Lingxiao Wang, and Quanquan Gu · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Later among the works it cites.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Closest in time.
How much over-parameterization is sufficient to learn deep relu networks?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuandong Tian · 2017
Cited alongside, same era.
Early stopping for kernel boosting algorithms: A general analysis with localized complexities
Yuting Wei, Fanny Yang, and Martin J Wainwright · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
Zixiang Chen, Yuan Cao, Difan Zou, and Quanquan Gu · 2019
Closest in time.
Robust statistical learning with lipschitz and convex loss functions
Geoffrey Chinot, Guillaume Lecué, and Matthieu Lerasle · 2019
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Closest in time.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Closest in time.
Global convergence of adaptive gradient methods for an over-parameterized neural network
Xiaoxia Wu, Simon S Du, and Rachel Ward · 2019
Closest in time.
Fast convergence of natural gradient descent for overparameterized neural networks
Guodong Zhang, James Martens, and Roger Grosse · 2019
Closest in time.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Closest in time.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks
Ziwei Ji and Matus Telgarsky · 2020
Closest in time.