Fetching the paper…
Reading the bibliography…
A recent line of research on deep learning focuses on the extremely over-parameterized setting, and shows that when the network width is larger than a high degree polynomial of the training sample size $n$ and the inverse of the target error $\epsilon^{-1}$, deep neural networks learned by (stochastic) gradient descent enjoy nice optimization and generalization guarantees.
Oymak, S · 1902
Earlier work this paper cites.
Nitanda, A · 1905
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Bartlett, P. L · 2002
Earlier work this paper cites.
Gradient methods never overfit on separable data
Shamir, O · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, R · 2010
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K · 2015
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Cited alongside, same era.
Risk and parameter convergence of logistic regression
Ji, Z · 2018
Cited alongside, same era.
Foundations of machine learning
Mohri, M · 2018
Cited alongside, same era.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Bai, Y · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y · 2019
Cited alongside, same era.
Gradient descent finds global minima for generalizable deep neural networks of practical sizes
Kawaguchi, K · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J · 2019
Closest in time.
Gradient descent optimizes over-parameterized deep ReLU networks
Zou, D · 2019
Closest in time.
An improved analysis of training over-parameterized deep neural networks
Zou, D · 2019
Closest in time.
Generalization error bounds of gradient descent for learning over-parameterized deep relu networks
Cao, Y · 2020
Closest in time.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks
Ji, Z · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On lazy training in differentiable programming
Chizat, L · 2019
Cited alongside, same era.
Algorithm-dependent generalization bounds for overparameterized deep residual networks
Frei, S · 2019
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z
Cited in the paper.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z
Cited in the paper.
On the convergence rate of training recurrent neural networks
Allen-Zhu, Z
Cited in the paper.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S
Cited in the paper.
Closest in time.
On the generalization ability of on-line learning algorithms
Cesa-Bianchi, N · 2057
Closest in time.