Fetching the paper…
Reading the bibliography…
Despite their practical success, a theoretical understanding of the loss landscape of neural networks has proven challenging due to the high-dimensional, non-convex, and highly nonlinear structure of such models.
Training a 3-node neural network is np-complete
Avrim Blum and Ronald L Rivest · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
A nonlinear programming algorithm for solving semidefinite programs via low-rank factorization
Samuel Burer and Renato DC Monteiro · 2003
Earlier work this paper cites.
Local minima and convergence in low-rank semidefinite programming
Samuel Burer and Renato DC Monteiro · 2005
Earlier work this paper cites.
Spatiotemporal elements of macaque v1 receptive fields
Nicole C Rust, Odelia Schwartz, J Anthony Movshon, and Eero P Simoncelli · 2005
Earlier work this paper cites.
Fast linear algebra is stable
James Demmel, Ioana Dumitriu, and Olga Holtz · 2007
Earlier work this paper cites.
Mnist handwritten digit database. at&t labs, 2010
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Global convergence of stochastic gradient descent for some non-convex matrix problems
Christopher De Sa, Kunle Olukotun, and Christopher Ré · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Cited alongside, same era.
The non-convex burer-monteiro approach works on smooth semidefinite programs
Nicolas Boumal, Vlad Voroninski, and Afonso Bandeira · 2016
Cited alongside, same era.
Global optimality of local search for low rank matrix recovery
Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Cited alongside, same era.
The expressive power of neural networks: A view from the width
Polynomial regression as an alternative to neural nets
Xi Cheng, Bohdan Khomtchouk, Norman Matloff, and Pete Mohanty · 2018
Later among the works it cites.
Universal approximation with quadratic deep networks
Fenglei Fan and Ge Wang · 2018
Later among the works it cites.
On the connection between learning two-layers neural networks and tensor decomposition
Marco Mondelli and Andrea Montanari · 2018
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
Simon Du and Jason Lee · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Cited alongside, same era.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Cited alongside, same era.
Convergence results for neural networks via electrodynamics
Rina Panigrahy, Ali Rahimi, Sushant Sachdeva, and Qiuyi Zhang · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Cited alongside, same era.
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2019
Closest in time.
On the expressive power of deep polynomial neural networks
Joe Kileel, Matthew Trager, and Joan Bruna · 2019
Closest in time.
Analysis of a two-layer neural network via displacement convexity
Adel Javanmard, Marco Mondelli, and Andrea Montanari · 2019
Closest in time.
Spurious valleys in one-hidden-layer neural network optimization landscapes
Luca Venturi, Afonso S Bandeira, and Joan Bruna · 2019
Closest in time.
Depth with nonlinearity creates no bad local minima in resnets
Kenji Kawaguchi and Yoshua Bengio · 2019
Closest in time.
Effect of depth and width on local minima in deep learning
Kenji Kawaguchi, Jiaoyang Huang, and Leslie Pack Kaelbling · 2019
Closest in time.
Stationary points of shallow neural networks with quadratic activation function
David Gamarnik, Eren C Kızıldağ, and Ilias Zadik · 2019
Closest in time.
Statistical mechanics of deep learning
Yasaman Bahri, Jonathan Kadmon, Jeffrey Pennington, Sam S Schoenholz, Jascha Sohl-Dickstein, and Surya Ganguli · 2020
Closest in time.