Fetching the paper…
Reading the bibliography…
We show that for any convex differentiable loss, a deep linear network has no spurious local minima as long as it is true for the two layer case.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
T. Zhang · 2004
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
Identiy matters in deep learning
M. Hardt and T. Ma · 2017
Earlier work this paper cites.
Depth creates no bad local minima
H. Lu and K. Kawaguchi · 2017
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Z. Allen-Zhu, Y. Li, and Y. Liang · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2018
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Cited alongside, same era.
Gradient descent with identity initialization efficiently learns positive definite linear transformations
P. Bartlett, D. Helmbold, and P. Long · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
S. S. Du, J. D. Lee, L. W. Haochuan Li and, and X. Zhai · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
S. S. Du, J. D. Lee, H. Li, L. Wang, and X. Zhai · 2018
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
T. Laurent and J. von Brecht · 2018
Cited alongside, same era.
Spurious valleys in two-layer neural network optimization landscapes
L. Venturi, A. S. Bandeira, and J. Bruna · 2018
Later among the works it cites.
Global optimality conditions for deep neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2018
Later among the works it cites.
Critical points of linear neural networks: Analytical forms and landscape properties
Y. Zhou and Y. Liang · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2018
Later among the works it cites.
A convergence analysis of gradient descent for deep linear neural networks
S. Arora, N. Cohen, N. Golowich, and W. Hu · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2018
Cited alongside, same era.
Small nonlinearities in activation functions create bad local minima in neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2019
Closest in time.