Fetching the paper…
Reading the bibliography…
Empirical studies show that gradient-based methods can learn deep neural networks (DNNs) with very good generalization performance in the over-parameterization regime, where DNNs can easily fit a random labeling of the training data.
Yang, G · 1902
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Bartlett, P. L · 2002
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Rahimi, A · 2009
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A · 2012
Earlier work this paper cites.
Learning without concentration
Mendelson, S · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S · 2014
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Haeffele, B. D · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Neyshabur, B · 2015
Earlier work this paper cites.
Representation benefits of deep feedforward networks
Telgarsky, M · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kawaguchi, K · 2016
Earlier work this paper cites.
On the quality of the initial basin in overspecified neural networks
Safran, I · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, D · 2016
Earlier work this paper cites.
benefits of depth in neural networks
Telgarsky, M · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L · 2017
Earlier work this paper cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Brutzkus, A · 2017
Earlier work this paper cites.
Sgd learns the conjugate kernel class of the network
Daniely, A · 2017
Earlier work this paper cites.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Dziugaite, G. K · 2017
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
Freeman, C. D · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Gunasekar, S · 2017
Cited alongside, same era.
Universal function approximation by deep neural nets with bounded width and relu activations
Hanin, B · 2017
Cited alongside, same era.
Approximating continuous functions by ReLU nets of minimal width
Hanin, B · 2017
Cited alongside, same era.
Identity matters in deep learning
Hardt, M · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Li, Y · 2017
Cited alongside, same era.
Why deep neural networks for function approximation?
Liang, S · 2017
ResNet with one-neuron hidden layers is a universal approximator
Lin, H · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
Mei, S · 2018
Later among the works it cites.
Foundations of machine learning
Mohri, M · 2018
Later among the works it cites.
Rotskoff, G. M · 2018
Later among the works it cites.
Spurious local minima are common in two-layer relu neural networks
Safran, I · 2018
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The expressive power of neural networks: A view from the width
Lu, Z · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Nguyen, Q · 2017
Cited alongside, same era.
Learning ReLUs via gradient descent
Soltanolkotabi, M · 2017
Cited alongside, same era.
Tian, Y · 2017
Cited alongside, same era.
Diverse neural network learns true target functions
Xie, B · 2017
Cited alongside, same era.
Error bounds for approximations with deep ReLU networks
Yarotsky, D · 2017
Cited alongside, same era.
Later among the works it cites.
A mean field view of the landscape of two-layers neural networks
Song, M · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D · 2018
Later among the works it cites.
Neural networks with finite intrinsic dimension have no spurious valleys
Venturi, L · 2018
Later among the works it cites.
Optimal approximation of continuous functions by very deep relu networks
Yarotsky, D · 2018
Later among the works it cites.
Global optimality conditions for deep neural networks
Yun, C · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D · 2018
Later among the works it cites.
Learning one-hidden-layer neural networks under general input distributions
Gao, W · 2019
Closest in time.
The implicit bias of gradient descent on nonseparable data
Ji, Z · 2019
Closest in time.
Just interpolate: Kernel” ridgeless” regression can generalize
Liang, T · 2019
Closest in time.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Nacson, M. S · 2019
Closest in time.
Mean field analysis of neural networks: A central limit theorem
Sirignano, J · 2019
Closest in time.
Regularization matters: Generalization and optimization of neural nets v.s. their induced kernel
Wei, C · 2019
Closest in time.
Small nonlinearities in activation functions create bad local minima in neural networks
Yun, C · 2019
Closest in time.
Learning one-hidden-layer ReLU networks via gradient descent
Zhang, X · 2019
Closest in time.
Gradient descent optimizes over-parameterized deep ReLU networks
Zou, D · 2019
Closest in time.
An improved analysis of training over-parameterized deep neural networks
Zou, D · 2019
Closest in time.