Fetching the paper…
Reading the bibliography…
We prove that for an $L$-layer fully-connected linear neural network, if the width of every hidden layer is $\tilde\Omega (L \cdot r \cdot d_{\mathrm{out}} \cdot \kappa^3 )$, where $r$ and $\kappa$ are the rank and the condition number of the input data, and $d_{\mathrm{out}}$ is the output dimension, then gradient descent with Gaussian random initialization converges to a global minimum at a linear rate.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J · 2016
Earlier work this paper cites.
Identity matters in deep learning
Hardt, M. and Ma, T · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kawaguchi, K · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, D. and Carmon, Y · 2016
Earlier work this paper cites.
Globally optimal gradient descent for a ConvNet with gaussian inputs
Brutzkus, A. and Globerson, A · 2017
Earlier work this paper cites.
Global optimality in neural network training
Haeffele, B. and Vidal, R · 2017
Earlier work this paper cites.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with ReLU activation
Li, Y. and Yuan, Y · 2017
Cited alongside, same era.
Depth creates no bad local minima
Lu, H. and Kawaguchi, K · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Nguyen, Q. and Hein, M · 2017
Cited alongside, same era.
Learning ReLUs via gradient descent
Soltanolkotabi, M · 2017
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Later among the works it cites.
Spurious local minima are common in two-layer relu neural networks
Safran, I. and Shamir, O · 2018
Later among the works it cites.
Exponential convergence time of gradient descent for one-dimensional deep linear neural networks
Shamir, O · 2018
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D · 2018
Later among the works it cites.
Neural networks with finite intrinsic dimension have no spurious valleys
Venturi, L., Bandeira, A., and Bruna, J · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tian, Y · 2017
Cited alongside, same era.
Global optimality conditions for deep neural networks
Yun, C., Sra, S., and Jadbabaie, A · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2018
Cited alongside, same era.
Gradient descent with identity initialization efficiently learns positive definite linear transformations
Bartlett, P., Helmbold, D., and Long, P · 2018
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
Du, S. S. and Lee, J. D · 2018
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
Laurent, T. and Brecht, J · 2018
Cited alongside, same era.
Zhang, X., Yu, Y., Wang, L., and Gu, Q · 2018
Later among the works it cites.
Critical points of linear neural networks: Analytical forms and landscape properties
Zhou, Y. and Liang, Y · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2019
Closest in time.
Gradient descent aligns the layers of deep linear networks
Ji, Z. and Telgarsky, M · 2019
Closest in time.
Small nonlinearities in activation functions create bad local minima in neural networks
Yun, C., Sra, S., and Jadbabaie, A · 2019
Closest in time.