Fetching the paper…
Reading the bibliography…
With the increasing popularity of non-convex deep models, developing a unifying theory for studying the optimization problems that arise from training these models becomes very significant.
Functional analysis, mcgraw-hill series in higher mathematics
W. Rudin · 1973
Earlier work this paper cites.
An elementary counterexample to the open mapping principle for bilinear maps
C. Horowitz · 1975
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Training a 3-node neural network is np-complete
A. Blum and R. L. Rivest · 1989
Earlier work this paper cites.
Weighted low-rank approximations
N. Srebro and T. Jaakkola · 2003
Earlier work this paper cites.
Local minima and convergence in low-rank semidefinite programming
S. Burer and R. D. C. Monteiro · 2005
Earlier work this paper cites.
Openness of multiplication in some function spaces
M. Balcerzak, A. Majchrzycki, and A. Wachowicz · 2013
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Earlier work this paper cites.
Matrix completion via nonconvex factorization: Algorithms and theory
R. Sun · 2015
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2016
Earlier work this paper cites.
The non-convex burer-monteiro approach works on smooth semidefinite programs
N. Boumal, V. Voroninski, and A. Bandeira · 2016
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
R. Ge, J. D. Lee, and T. Ma · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
Non-square matrix sensing without spurious local minima via the burer-monteiro approach
D. Park, A. Kyrillidis, C. Caramanis, and S. Sanghavi · 2016
Cited alongside, same era.
A unified computational and statistical framework for nonconvex low-rank matrix estimation
L. Wang, X. Zhang, and Q. Gu · 2016
Cited alongside, same era.
Q. Zheng and J. Lafferty · 2016
Cited alongside, same era.
Where is matrix multiplication locally open?
E. Behrends · 2017
Cited alongside, same era.
Depth creates no bad local minima
H. Lu and K. Kawaguchi · 2017
On the loss landscape of a class of deep neural networks with no bad local valleys
Quynh Nguyen, Mahesh Chandra Mukkamala, and Matthias Hein · 2018
Closest in time.
Spurious valleys in two-layer neural network optimization landscapes
Luca Venturi, Afonso S Bandeira, and Joan Bruna · 2018
Closest in time.
Small nonlinearities in activation functions create bad local minima in neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Cited alongside, same era.
Global optimality conditions for deep neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2017
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2018
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
Simon S Du and Jason D Lee · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Closest in time.
Sub-optimal local minima exist for almost all over-parameterized neural networks
Tian Ding, Dawei Li, and Ruoyu Sun · 2019
Closest in time.
On connected sublevel sets in deep learning
Quynh Nguyen · 2019
Closest in time.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2019
Closest in time.
Depth creates no more spurious local minima
Li Zhang · 2019
Closest in time.
Distributed low-rank matrix factorization with exact consensus
Zhihui Zhu, Qiuwei Li, Xinshuo Yang, Gongguo Tang, and Michael B Wakin · 2019
Closest in time.
The global optimization geometry of shallow linear neural networks
Zhihui Zhu, Daniel Soudry, Yonina C Eldar, and Michael B Wakin · 2019
Closest in time.
The global geometry of centralized and distributed low-rank matrix recovery without regularization
Shuang Li, Qiuwei Li, Zhihui Zhu, Gongguo Tang, and Michael B Wakin · 2020
Closest in time.
Global convergence of maml for lqr
Igor Molybog and Javad Lavaei · 2020
Closest in time.