Fetching the paper…
Reading the bibliography…
We study the stability and convergence of training deep ResNets with gradient descent.
Over parameterized two-level neural networks can learn near optimal feature representations
Cong Fang, Hanze Dong, and Tong Zhang · 1910
Earlier work this paper cites.
Convex formulation of overparameterized deep neural networks
Cong Fang, Yihong Gu, Weizhong Zhang, and Tong Zhang · 1911
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Smoothed analysis of algorithms: Why the simplex algorithm usually takes polynomial time
Daniel A Spielman and Shang-Hua Teng · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
R. Vershynin · 2012
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
Andreas Veit, Michael J Wilber, and Serge Belongie · 2016
Earlier work this paper cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
David Balduzzi, Marcus Frean, Lennox Leary, JP Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Earlier work this paper cites.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Earlier work this paper cites.
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Mean field residual networks: On the edge of chaos
Ge Yang and Samuel Schoenholz · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
A generalization theory of gradient descent for learning over-parameterized deep ReLU networks
Yuan Cao and Quanquan Gu · 2019
Closest in time.
How much over-parameterization is sufficient to learn deep ReLU networks?
Zixiang Chen, Yuan Cao, Difan Zou, and Quanquan Gu · 2019
Closest in time.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Closest in time.
Algorithm-dependent generalization bounds for overparameterized deep residual networks
Spencer Frei, Yuan Cao, and Quanquan Gu · 2019
Closest in time.
Limitations of lazy training of two-layers neural networks
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
Skip connections eliminate singularities
A Emin Orhan and Xaq Pitkow · 2018
Cited alongside, same era.
On the local hessian in back-propagation
Huishuai Zhang, Wei Chen, and Tie-Yan Liu · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Cited alongside, same era.
What can ResNet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Cited alongside, same era.
Ziwei Ji and Matus Telgarsky · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Closest in time.
The role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2019
Closest in time.
Mean field limit of the learning dynamics of multilayer neural networks
Phan-Minh Nguyen · 2019
Closest in time.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2019
Closest in time.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Closest in time.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Closest in time.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2020
Closest in time.