Fetching the paper…
Reading the bibliography…
The skip-connections used in residual networks have become a standard architecture choice in deep learning due to the increased training stability and generalization performance with this architecture, although there has been limited theoretical understanding for this improvement.
Generalization error bounds of gradient descent for learning over-parameterized deep relu networks
Cao, Y · 1902
Earlier work this paper cites.
Training over-parameterized deep resnet is almost as easy as training a two-layer network
Zhang, H · 1903
Earlier work this paper cites.
Temporal convolution for real-time keyword spotting on mobile devices
Choi, S · 1904
Earlier work this paper cites.
Analysis of the gradient descent algorithm for a deep neural network model with skip-connections
E, W · 1904
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Rahimi, A · 2008
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, R · 2010
Earlier work this paper cites.
Understanding Machine Learning: From Theory to Algorithms
Shalev-Shwartz, S · 2014
Earlier work this paper cites.
Convolutional neural networks for small-footprint keyword spotting
Sainath, T. N · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K · 2016
Earlier work this paper cites.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and <1mb model size
Iandola, F. N · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L · 2017
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Dziugaite, G. K · 2017
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A · 2017
Cited alongside, same era.
Error bounds for approximations with deep relu networks
Yarotsky, D · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C · 2017
On tighter generalization bound for deep neural networks: Cnns, resnets, and beyond
Li, X · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y · 2018
Later among the works it cites.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Neyshabur, B · 2018
Later among the works it cites.
Deep residual learning for small-footprint keyword spotting
Tang, R · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z · 2019
Closest in time.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Arora, S · 2018
Cited alongside, same era.
Size-independent sample complexity of neural networks
Golowich, N · 2018
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y
Cited in the paper.
Gradient descent finds global minima of deep neural networks
Du, S. S
Cited in the paper.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S
Cited in the paper.
Arora, S · 2019
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D · 2019
Closest in time.
An improved analysis of training over-parameterized deep neural networks
Zou, D · 2019
Closest in time.