Fetching the paper…
Reading the bibliography…
A recent line of research has shown that gradient-based algorithms with random initialization can converge to the global minima of the training loss for over-parameterized (i.e., sufficiently wide) deep neural networks.
Arora, S · 1901
Earlier work this paper cites.
A generalization theory of gradient descent for learning over-parameterized deep relu networks
Cao, Y · 1902
Earlier work this paper cites.
Oymak, S · 1902
Earlier work this paper cites.
Global convergence of adaptive gradient methods for an over-parameterized neural network
Wu, X · 1902
Earlier work this paper cites.
Training over-parameterized deep resnet is almost as easy as training a two-layer network
Zhang, H · 1903
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C · 2016
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Brutzkus, A · 2017
Cited alongside, same era.
When is a convolutional filter easy to learn?
Du, S. S · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with ReLU activation
Li, Y · 2017
Cited alongside, same era.
Tian, Y · 2017
Cited alongside, same era.
A note on lazy training in supervised differentiable programming
Chizat, L · 2018
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
Du, S. S · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y · 2018
Later among the works it cites.
Learning one-hidden-layer ReLU networks via gradient descent
Zhang, X · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Recovery guarantees for one-hidden-layer neural networks
Zhong, K · 2017
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z
Cited in the paper.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z
Cited in the paper.
On the convergence rate of training recurrent neural networks
Allen-Zhu, Z
Cited in the paper.
Gradient descent finds global minima of deep neural networks
Du, S. S
Cited in the paper.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S
Cited in the paper.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D · 2018
Later among the works it cites.