Fetching the paper…
Reading the bibliography…
Traditional landscape analysis of deep neural networks aims to show that no sub-optimal local minima exist in some appropriate sense.
Interference alignment using finite and dependent channel extensions: The single beam case
Ruoyu Sun and Zhi-Quan Luo · 1912
Earlier work this paper cites.
Theory of the integral
S. Saks · 1937
Earlier work this paper cites.
Nonlinear Programming, 2nd ed
D. P. Bertsekas · 1999
Earlier work this paper cites.
Convergence of a block coordinate descent method for nondifferentiable minimization
P. Tseng · 2001
Earlier work this paper cites.
Matrix completion from a few entries
Raghunandan H Keshavan, Andrea Montanari, and Sewoong Oh · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Perturbation theory for linear operators
T. Kato · 2013
Earlier work this paper cites.
Guaranteed matrix completion via non-convex factorization
R. Sun and Z.-Q. Luo · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Earlier work this paper cites.
Phase retrieval via wirtinger flow: Theory and algorithms
Emmanuel J Candes, Xiaodong Li, and Mahdi Soltanolkotabi · 2015
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C D. Freeman and J. Bruna · 2016
Earlier work this paper cites.
Gradient descent converges to minimizers
J. D Lee, M. Simchowitz, M. I Jordan, and B. Recht · 2016
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro · 2017
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
P. L. Bartlett, D. J. Foster, and M. J. Telgarsky · 2017
Earlier work this paper cites.
Y. Li, T. Ma, and H. Zhang · 2017
Cited alongside, same era.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
D. Soudry and E. Hoffer · 2017
Cited alongside, same era.
Global optimality in neural network training
Benjamin D Haeffele and René Vidal · 2017
Cited alongside, same era.
Learning one-hidden-layer neural networks with landscape design
R. Ge, J. D Lee, and T. Ma · 2017
Cited alongside, same era.
Porcupine neural networks:(almost) all local optima are global
Soheil Feizi, Hamid Javadi, Jesse Zhang, and David Tse · 2017
On the convergence of adaptive gradient methods for nonconvex optimization
D. Zhou, Y. Tang, Z. Yang, Y. Cao, and Q. Gu · 2018
Later among the works it cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Later among the works it cites.
On the margin theory of feedforward neural networks
C. Wei, J. D Lee, Q. Liu, and T. Ma · 2018
Later among the works it cites.
Towards understanding the role of over-parametrization in generalization of neural networks
B. Neyshabur, Z. Li, S. Bhojanapalli, Y. LeCun, and N. Srebro · 2018
Later among the works it cites.
A mean field view of the landscape of two-layers neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Convergence results for neural networks via electrodynamics
Rina Panigrahy, Ali Rahimi, Sushant Sachdeva, and Qiuyi Zhang · 2017
Cited alongside, same era.
The multilinear structure of relu networks
Thomas Laurent and James von Brecht · 2017
Cited alongside, same era.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
A. Brutzkus and A. Globerson · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
K. Zhong, Z. Song, P. Jain, P. L Bartlett, and I. S Dhillon · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Y. Li and Y. Yuan · 2017
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Cited alongside, same era.
S. Mei, A. Montanari, and P. Nguyen · 2018
Later among the works it cites.
Mean field analysis of neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Later among the works it cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Later among the works it cites.
Grant M Rotskoff and Eric Vanden-Eijnden · 2018
Later among the works it cites.
Learning one-hidden-layer neural networks under general input distributions
Weihao Gao, Ashok Vardhan Makkuva, Sewoong Oh, and Pramod Viswanath · 2018
Later among the works it cites.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz · 2018
Later among the works it cites.
Learning relu networks on linearly separable data: Algorithm, optimality, and generalization
Gang Wang, Georgios B Giannakis, and Jie Chen · 2018
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
S. S Du and J. D Lee · 2018
Later among the works it cites.
On the connection between learning two-layers neural networks and tensor decomposition
Marco Mondelli and Andrea Montanari · 2018
Later among the works it cites.
On the convergence of adagrad with momentum for training deep neural networks
Fangyu Zou and Li Shen · 2018
Later among the works it cites.
Gradient descent with random initialization: Fast global convergence for nonconvex phase retrieval
Yuxin Chen, Yuejie Chi, Jianqing Fan, and Cong Ma · 2018
Later among the works it cites.
A geometric analysis of phase retrieval
Ju Sun, Qing Qu, and John Wright · 2018
Later among the works it cites.
Elimination of all bad local minima in deep learning
K. Kawaguchi and L. P. Kaelbling · 2019
Closest in time.
Spurious local minima exist for almost all over-parameterized neural networks
Tian Ding, Dawei Li, and Ruoyu Sun · 2019
Closest in time.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2019
Closest in time.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Closest in time.