Fetching the paper…
Reading the bibliography…
We study the convergence of gradient flows related to learning deep linear neural networks (where the activation function is the identity map) from data.
Sur les trajectoires du gradient d’une fonction analytique
S. Lojasiewicz · 1984
Earlier work this paper cites.
Global Stability of Dynamical Systems
M. Shub · 1986
Earlier work this paper cites.
Global analysis of Oja’s flow for neural networks
W. Yan, U. Helmke, and J. B. Moore · 1994
Earlier work this paper cites.
Critical points of matrix least squares distance functions
U. Helmke and M. A. Shayman · 1995
Earlier work this paper cites.
Theorems on Regularity and Singularity of Energy Minimizing Maps
L. Simon · 1996
Earlier work this paper cites.
Matrix Analysis
R. Bhatia · 1997
Earlier work this paper cites.
Fundamentals of Differential Geometry
S. Lang · 1999
Earlier work this paper cites.
Proof of the gradient conjecture of R. Thom
K. Kurdyka, T. Mostowski, and A. Parusinski · 2000
Earlier work this paper cites.
Lectures on Morse homology
A. Banyaga, D. Hurtubise · 2004
Cited alongside, same era.
Convergence of the iterates of descent methods for analytic cost functions
P. A. Absil, R. Mahony, and B. Andrews · 2005
Cited alongside, same era.
Eulerian calculus for the contraction in the Wasserstein distance
F. Otto and M. Westdickenberg · 2005
Cited alongside, same era.
Optimization Algorithms on Matrix Manifolds
P.-A. Absil, R. Mahony, and R. Sepulchre · 2008
Cited alongside, same era.
The operator equation ∑ i = 0 n A n − i X B i = Y \sum_{i=0}^{n}A^{n-i}XB^{i}=Y
R. Bhatia, M. Uchiyama · 2009
Cited alongside, same era.
An analog of the 2-Wasserstein metric in non-commutative probability under which the fermionic Fokker-Planck equation is gradient flow for the entropy
E. Carlen and J. Maas · 2012
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Later among the works it cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Later among the works it cites.
A geometric approach of gradient descent algorithms in neural networks
Y. Chitour, Z. Liao, and R. Couillet · 2018
Later among the works it cites.
S. Arora, N. Cohen, N. Golowich, and W. Hu · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Machine learning: a probabilistic perspective
K. P. Murphy · 2013
Cited alongside, same era.
Low-rank retractions: a survey and new results
P.-A. Absil, I.V. Oseledets · 2015
Cited alongside, same era.
First-order methods almost always avoid strict saddle points
J. D. Lee, I. Panageas, G. Piliouras, M. Simchowitz, M. I. Jordan, and B. Recht · 2019
Closest in time.
Pure and spurious critical points: a geometric study of linear networks
M. Trager, K. Kohn, J. Bruna, · 2019
Closest in time.