Fetching the paper…
Reading the bibliography…
In this paper, we propose a geometric framework to analyze the convergence properties of gradient descent trajectories in the context of linear neural networks.
Sur les trajectoires du gradient d’une fonction analytique
S Lojasiewicz · 1982
Earlier work this paper cites.
Some NP-complete problems in quadratic and nonlinear programming
Katta G Murty and Santosh N Kabadi · 1987
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
Avrim Blum and Ronald L Rivest · 1989
Earlier work this paper cites.
Matrix analysis
Roger A Horn and Charles R Johnson · 1990
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Lectures on partial hyperbolicity and stable ergodicity
Ya B Pesin and Yakov B Pesin · 2004
Earlier work this paper cites.
Convergence of the iterates of descent methods for analytic cost functions
Pierre-Antoine Absil, Robert Mahony, and Benjamin Andrews · 2005
Earlier work this paper cites.
Explicit bounds for the Łojasiewicz exponent in the gradient inequality for polynomials
Didier D’Acunto and Krzysztof Kurdyka · 2005
Earlier work this paper cites.
Invariant manifolds
Morris W Hirsch, Charles Chapman Pugh, and Michael Shub · 2006
Earlier work this paper cites.
Nonsmooth analysis and control theory
Francis H Clarke, Yuri S Ledyaev, Ronald J Stern, and Peter R Wolenski · 2008
Cited alongside, same era.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Ronan Collobert and Jason Weston · 2008
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Acoustic modeling using deep belief networks
Abdel-rahman Mohamed, George E Dahl, and Geoffrey Hinton · 2012
Cited alongside, same era.
Ordinary differential equations and dynamical systems
Gerald Teschl · 2012
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Later among the works it cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Later among the works it cites.
Gradient descent only converges to minimizers: Non-isolated critical points and invariant regions
Ioannis Panageas and Georgios Piliouras · 2017
Later among the works it cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Closest in time.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Regression modeling strategies: with applications to linear models, logistic and ordinal regression, and survival analysis
Frank E Harrell Jr · 2015
Cited alongside, same era.
When are nonconvex problems not scary?
Ju Sun, Qing Qu, and John Wright · 2015
Cited alongside, same era.
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Closest in time.
Deep linear networks with arbitrary loss: All local minima are global
Thomas Laurent and James Brecht · 2018
Closest in time.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Closest in time.
First-order methods almost always avoid strict saddle points
Jason D. Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2019
Closest in time.
Convergence conditions for nonlinear programming algorithms
Willard I Zangwill · 2022
Closest in time.