Fetching the paper…
Reading the bibliography…
We study the training and generalization of deep neural networks (DNNs) in the over-parameterized regime, where the network width (i.e., number of hidden nodes per layer) is much larger than the number of training data points.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J · 1902
Earlier work this paper cites.
Oymak, S · 1902
Earlier work this paper cites.
Yang, G · 1902
Earlier work this paper cites.
E, W · 1904
Earlier work this paper cites.
On the power and limitations of random features for understanding neural networks
Yehudai, G · 1904
Earlier work this paper cites.
Size-free generalization bounds for convolutional neural networks
Long, P. M · 1905
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y · 1998
Earlier work this paper cites.
(not) bounding the true error
Langford, J · 2002
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A · 2008
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Rahimi, A · 2009
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A · 2012
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S · 2014
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K · 2015
Cited alongside, same era.
Norm-based capacity control in neural networks
Neyshabur, B · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L · 2017
Cited alongside, same era.
Size-independent sample complexity of neural networks
Golowich, N · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Later among the works it cites.
On tighter generalization bound for deep neural networks: Cnns, resnets, and beyond
Li, X · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y · 2018
Later among the works it cites.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Neyshabur, B · 2018
Later among the works it cites.
Making the last iterate of sgd information theoretically optimal
Jain, P · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sgd learns the conjugate kernel class of the network
Daniely, A · 2017
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Dziugaite, G. K · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C · 2017
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Arora, S · 2018
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z
Cited in the paper.
Closest in time.
Towards understanding the role of over-parametrization in generalization of neural networks
Neyshabur, B · 2019
Closest in time.
Regularization matters: Generalization and optimization of neural nets v.s. their induced kernel
Wei, C · 2019
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Zou, D · 2019
Closest in time.
Generalization error bounds of gradient descent for learning over-parameterized deep relu networks
Cao, Y · 2020
Closest in time.
On the generalization ability of on-line learning algorithms
Cesa-Bianchi, N · 2057
Closest in time.