Fetching the paper…
Reading the bibliography…
One of the mysteries in the success of neural networks is randomly initialized first order methods like gradient descent can achieve zero training loss even though the objective function is non-convex and non-smooth.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Earlier work this paper cites.
Escaping from saddle points − - online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Benjamin D Haeffele and René Vidal · 2015
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
On the quality of the initial basin in overspecified neural networks
Itay Safran and Ohad Shamir · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Earlier work this paper cites.
A Lyapunov analysis of momentum methods in optimization
Ashia C Wilson, Benjamin Recht, and Michael I Jordan · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Globally optimal gradient descent for a ConvNet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Earlier work this paper cites.
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Cited alongside, same era.
Learning ReLus via gradient descent
Mahdi Soltanolkotabi · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Closest in time.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Closest in time.
Stochastic subgradient method converges on tame functions
Damek Davis, Dmitriy Drusvyatskiy, Sham Kakade, and Jason D Lee · 2018
Closest in time.
On the power of over-parametrization in neural networks with quadratic activation
Simon S Du and Jason D Lee · 2018
Closest in time.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Cited alongside, same era.
Invariance of weight distributions in rectified mlps
Russell Tsuchida, Farbod Roosta-Khorasani, and Marcus Gallagher · 2017
Cited alongside, same era.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2017
Cited alongside, same era.
Critical points of neural networks: Analytical forms and landscape properties
Yi Zhou and Yingbin Liang · 2017
Cited alongside, same era.
Gradient descent can take exponential time to escape saddle points
Simon S Du, Chi Jin, Jason D Lee, Michael I Jordan, Aarti Singh, and Barnabas Poczos
Cited in the paper.
When is a convolutional filter easy to learn?
Simon S Du, Jason D Lee, and Yuandong Tian
Cited in the paper.
Closest in time.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Closest in time.
Optimization landscape and expressivity of deep cnns
Quynh Nguyen and Matthias Hein · 2018
Closest in time.
Spurious local minima are common in two-layer ReLU neural networks
Itay Safran and Ohad Shamir · 2018
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2018
Closest in time.
Neural networks with finite intrinsic dimension have no spurious valleys
Luca Venturi, Afonso Bandeira, and Joan Bruna · 2018
Closest in time.