Fetching the paper…
Reading the bibliography…
We study the optimization landscape of deep linear neural networks with the square loss.
Some np-complete problems in quadratic and nonlinear programming
Katta G Murty and Santosh N Kabadi · 1987
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Training a 3-node neural network is np-complete
Avrim Blum and Ronald L Rivest · 1989
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Yurii Nesterov · 1998
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural network
Andrew M Saxe, James L Mcclelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Earlier work this paper cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
Earlier work this paper cites.
Depth creates no bad local minima
Haihao Lu and Kenji Kawaguchi · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Earlier work this paper cites.
Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks
Peter Bartlett, Dave Helmbold, and Philip Long · 2018
Earlier work this paper cites.
Escaping saddles with stochastic gradients
Hadi Daneshmand, Jonas Kohler, Aurelien Lucchi, and Thomas Hofmann · 2018
Earlier work this paper cites.
Accelerated gradient descent escapes saddle points faster than gradient descent
Chi Jin, Praneeth Netrapalli, and Michael I Jordan · 2018
Earlier work this paper cites.
Deep linear networks with arbitrary loss: All local minima are global
Thomas Laurent and James von Brecht · 2018
Earlier work this paper cites.
Global optimality conditions for deep neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Earlier work this paper cites.
Critical points of linear neural networks: Analytical forms and landscape properties
Yi Zhou and Yingbin Liang · 2018
Cited alongside, same era.
Local saddle point optimization: A curvature exploitation approach
Leonard Adolphs, Hadi Daneshmand, Aurelien Lucchi, and Thomas Hofmann · 2019
Cited alongside, same era.
Width provably matters in optimization for deep linear neural networks
Simon Du and Wei Hu · 2019
Cited alongside, same era.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Cited alongside, same era.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2019
Cited alongside, same era.
First-order methods almost always avoid strict saddle points
Jason D Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2019
Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
Mikhail Belkin · 2021
Closest in time.
On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M. Kakade, and Michael I. Jordan · 2021
Closest in time.
The loss surface of deep linear networks viewed through the algebraic geometry lens
Dhagash Mehta, Tianran Chen, Tingting Tang, and Jonathan Hauenstein · 2021
Closest in time.
Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers
Bubacarr Bah, Holger Rauhut, Ulrich Terstiege, and Michael Westdickenberg · 2022
Closest in time.
Asymptotic study of stochastic adaptive algorithm in non-convex landscape
Sébastien Gadat and Ioana Gavra · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A mathematical theory of semantic development in deep neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2019
Cited alongside, same era.
Exponential convergence time of gradient descent for one-dimensional deep linear neural networks
Ohad Shamir · 2019
Cited alongside, same era.
Global convergence of gradient descent for deep linear residual networks
Lei Wu, Qingcan Wang, and Chao Ma · 2019
Cited alongside, same era.
Training linear neural networks: Non-local convergence and complexity results
Armin Eftekhari · 2020
Cited alongside, same era.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Cited alongside, same era.
Pure and spurious critical points: a geometric study of linear networks
Matthew Trager, Kathlén Kohn, and Joan Bruna · 2020
Cited alongside, same era.
Arthur Jacot · 2022
Closest in time.
Arthur Jacot, François Ged, Berfin Şimşek, Clément Hongler, and Franck Gabriel · 2022
Closest in time.
Learning deep models: Critical points and local openness
Maher Nouiehed and Meisam Razaviyayn · 2022
Closest in time.
On the effective number of linear regions in shallow univariate relu networks: Convergence guarantees and implicit bias
Itay Safran, Gal Vardi, and Jason D Lee · 2022
Closest in time.
A geometric approach of gradient descent algorithms in linear neural networks
Yacine Chitour, Zhenyu Liao, and Romain Couillet · 2023
Closest in time.
A line-search descent algorithm for strict saddle functions with complexity guarantees
Michael J. O’Neill and Stephen J. Wright · 2023
Closest in time.
Implicit regularization towards rank minimization in relu networks
Nadav Timor, Gal Vardi, and Ohad Shamir · 2023
Closest in time.
Implicit regularization of deep residual networks towards neural odes
Pierre Marion, Yu-Han Wu, Michael E. Sander, and Gérard Biau · 2024
Closest in time.
Convergence of gradient descent for learning linear neural networks
Gabin Maxime Nguegnang, Holger Rauhut, and Ulrich Terstiege · 2024
Closest in time.