Fetching the paper…
Reading the bibliography…
We analyze the global convergence of gradient descent for deep linear residual networks by proposing a new initialization: zero-asymmetric (ZAS) initialization.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2017
Earlier work this paper cites.
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Earlier work this paper cites.
Fashion-mnist: A novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Gradient descent with identity initialization efficiently learns positive definite linear transformations
Peter Bartlett, Dave Helmbold, and Phil Long · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon S. Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2018
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2018
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
Thomas Laurent and James Brecht · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Closest in time.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2019
Closest in time.
Width provably matters in optimization for deep linear neural networks
Simon S. Du and Wei Hu · 2019
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A mean field view of the landscape of two-layers neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
Grant M. Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
Exponential convergence time of gradient descent for one-dimensional deep linear neural networks
Ohad Shamir · 2018
Cited alongside, same era.
Weinan E, Chao Ma, Qingcan Wang, and Lei Wu · 2019
Closest in time.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 2019
Closest in time.