Fetching the paper…
Reading the bibliography…
We study the convergence of gradient descent (GD) and stochastic gradient descent (SGD) for training $L$-hidden-layer linear residual networks (ResNets).
Inequalities for the trace of matrix product
Yuguang Fang, Kenneth A Loparo, and Xiangbo Feng · 1994
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
Daniel C Freeman and Joan Bruna · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with ReLU activation
Yuanzhi Li and Yang Yuan · 2017
Earlier work this paper cites.
Depth creates no bad local minima
Haihao Lu and Kenji Kawaguchi · 2017
Earlier work this paper cites.
An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad E Hazan · 2018
Earlier work this paper cites.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Exponential convergence time of gradient descent for one-dimensional deep linear neural networks
Ohad Shamir · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Later among the works it cites.
Algorithm-dependent generalization bounds for overparameterized deep residual networks
Spencer Frei, Yuan Cao, and Quanquan Gu · 2019
Later among the works it cites.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2019
Later among the works it cites.
Gradient descent finds global minima for generalizable deep neural networks of practical sizes
Kenji Kawaguchi and Jiaoyang Huang · 2019
Later among the works it cites.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Global optimality conditions for deep neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Cited alongside, same era.
Learning one-hidden-layer ReLU networks via gradient descent
Xiao Zhang, Yaodong Yu, Lingxiao Wang, and Quanquan Gu · 2018
Cited alongside, same era.
Critical points of linear neural networks: Analytical forms and landscape properties
Yi Zhou and Yingbin Liang · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Gradient descent with identity initialization efficiently learns positive-definite linear transformations by deep residual networks
Peter L Bartlett, David P Helmbold, and Philip M Long · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Width provably matters in optimization for deep linear neural networks
Simon Du and Wei Hu · 2019
Cited alongside, same era.
Lili Su and Pengkun Yang · 2019
Later among the works it cites.
Global convergence of gradient descent for deep linear residual networks
Lei Wu, Qingcan Wang, and Chao Ma · 2019
Later among the works it cites.
Training over-parameterized deep resnet is almost as easy as training a two-layer network
Huishuai Zhang, Da Yu, Wei Chen, and Tie-Yan Liu · 2019
Later among the works it cites.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2019
Later among the works it cites.
Provable benefit of orthogonal initialization in optimizing deep linear networks
Wei Hu, Lechao Xiao, and Jeffrey Pennington · 2020
Closest in time.