Backpropagation can give rise to spurious local minima even for networks without hidden layers
Eduardo D Sontag and Héctor J Sussmann · 1989
Earlier work this paper cites.
On the local minima free condition of backpropagation learning
Xiao-Hu Yu and Guo-An Chen · 1995
Earlier work this paper cites.
Exponentially many local minima for single neurons
Peter Auer, Mark Herbster, and Manfred K Warmuth · 1996
Earlier work this paper cites.
Optimal learning in artificial neural networks: A review of theoretical results
Monica Bianchini and Marco Gori · 1996
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Original
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Original
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Earlier work this paper cites.
The zero set of a real analytic function
Original
Boris Mityagin · 2015
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2016
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C D. Freeman and J. Bruna · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
Why deep neural networks for function approximation?
Original
Shiyu Liang and Rayadurgam Srikant · 2016
Earlier work this paper cites.
Guaranteed matrix completion via non-convex factorization
Ruoyu Sun and Zhi-Quan Luo · 2016
Earlier work this paper cites.
Local minima in training of deep networks
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2016
Earlier work this paper cites.
Benefits of depth in neural networks
Original
Matus Telgarsky · 2016
Earlier work this paper cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Original
A. Brutzkus and A. Globerson · 2017
Earlier work this paper cites.
Porcupine neural networks:(almost) all local optima are global
Original
Soheil Feizi, Hamid Javadi, Jesse Zhang, and David Tse · 2017
Earlier work this paper cites.
Global optimality in neural network training
Benjamin D Haeffele and René Vidal · 2017
Earlier work this paper cites.
The multilinear structure of relu networks
Original
Thomas Laurent and James von Brecht · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Y. Li and Y. Yuan · 2017
Earlier work this paper cites.
Depth creates no bad local minima
Original
Haihao Lu and Kenji Kawaguchi · 2017
Earlier work this paper cites.