Fetching the paper…
Reading the bibliography…
We study the convergence properties of gradient descent for training deep linear neural networks, i.e., deep matrix factorizations, by extending a previous analysis for the related gradient flow.
Convergence of the iterates of descent methods for analytic cost functions
P.-A. Absil, R. Mahony, and B. Andrews · 2005
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
Implicit regularization in matrix factorization
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Earlier work this paper cites.
Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks
P. Bartlett, D. Helmbold, and P. Long · 2018
Earlier work this paper cites.
A geometric approach of gradient descent algorithms in neural networks
Y. Chitour, Z. Liao, and R. Couillet · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
S. S. Du, W. Hu, and J. D. Lee · 2018
Earlier work this paper cites.
Global optimality conditions for deep neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2018
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
S. Arora, N. Cohen, N. Golowich, and W. Hu · 2019
Cited alongside, same era.
Implicit regularization in deep matrix factorization
S. Arora, N. Cohen, W. Hu, and Y. Luo · 2019
Cited alongside, same era.
Width provably matters in optimization for deep linear neural networks
S. Du and W. Hu · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2019
Cited alongside, same era.
First-order methods almost always avoid strict saddle points
J. D. Lee, I. Panageas, G. Piliouras, M. Simchowitz, M. I. Jordan, and B. Recht · 2019
Cited alongside, same era.
Gradient descent for deep matrix factorization: Dynamics and implicit bias towards low rank
H. Chou, C. Gieshoff, J. Maly, and H. Rauhut · 2020
Later among the works it cites.
Stochastic subgradient method converges on tame functions
D. Davis, D. Drusvyatskiy, S. Kakade, and J. D. Lee · 2020
Later among the works it cites.
Low-rank regularization and solution uniqueness in over-parameterized matrix sensing
K. Geyer, A. Kyrillidis, and A. Kalev · 2020
Later among the works it cites.
Provable benefit of orthogonal initialization in optimizing deep linear networks
W. Hu, L. Xiao, and J. Pennington · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
N. Razin and N. Cohen · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
First-order methods almost always avoid saddle points: The case of vanishing step-sizes
I. Panageas, G. Piliouras, and X. Wang · 2019
Cited alongside, same era.
Exponential convergence time of gradient descent for one-dimensional deep linear neural networks
O. Shamir · 2019
Cited alongside, same era.
Global convergence of gradient descent for deep linear residual networks
L. Wu, Q. Wang, and C. Ma · 2019
Cited alongside, same era.
Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers
B. Bah, H. Rauhut, U. Terstiege, and M. Westdickenberg
Cited in the paper.
Pure and spurious critical points: a geometric study of linear networks
M. Trager, K. Kohn, and J. Bruna · 2020
Later among the works it cites.
On the global convergence of training deep linear resnets
D. Zou, P. M. Long, and Q. Gu · 2020
Later among the works it cites.
Continuous vs. discrete optimization of deep neural networks
O. Elkabetz and N. Cohen · 2021
Closest in time.
A unifying view on implicit bias in training linear neural networks
C. Yun, S. Krishnan, and H. Mobahi · 2021
Closest in time.