Fetching the paper…
Reading the bibliography…
Recent works have shown that gradient descent can find a global minimum for over-parameterized neural networks where the widths of all the hidden layers scale polynomially with $N$ ($N$ being the number of training samples).
Global convergence of adaptive gradient methods for an over-parameterized neural network, 2019
Xiaoxia Wu, Simon S Du, and Rachel Ward · 1902
Earlier work this paper cites.
A. Nitanda, G. Chinot, and T. Suzuki · 1905
Earlier work this paper cites.
Quadratic suffices for over-parametrization via matrix chernoff bound, 2020
Z. Song and X. Yang · 1906
Earlier work this paper cites.
Mildly overparametrized neural nets can memorize training data efficiently, 2019
Rong Ge, Runzhe Wang, and Haoyu Zhao · 1909
Earlier work this paper cites.
How much over-parameterization is sufficient to learn deep ReLU networks?, 2019
Z. Chen, Y. Cao, D. Zou, and Q. Gu · 1911
Earlier work this paper cites.
Gradient methods for minimizing functionals
B. T. Polyak · 1963
Earlier work this paper cites.
On the capabilities of multilayer perceptrons
E. B. Baum · 1988
Earlier work this paper cites.
What size net gives valid generalization?
Eric B Baum and David Haussler · 1989
Earlier work this paper cites.
Perturbation theory for the singular value decomposition
G. W. Stewart · 1990
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
Avrim L Blum and Ronald L Rivest · 1992
Earlier work this paper cites.
Exponentially many local minima for single neurons
P. Auer, M. Herbster, and M. Warmuth · 1996
Earlier work this paper cites.
On the infinite width limit of neural networks with a standard parameterization, 2020
J. Sohl-Dickstein, R. Novak, S. S. Schoenholz, and J. Lee · 2001
Earlier work this paper cites.
Memory capacity of neural networks with threshold and ReLU activations, 2020
Roman Vershynin · 2001
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices, 2010
R. Vershynin · 2010
Earlier work this paper cites.
Restricted isometry property of matrices with independent columns and neighborly polytopes by random sampling
R. Adamczak, A. E. Litvak, A. Pajor, and N. Tomczak-Jaegermann · 2011
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
S. Boucheron, G. Lugosi, and P. Massart · 2013
Cited alongside, same era.
Sharp interpolation inequalities on the sphere: new methods and consequences
J. Dolbeault, M. J. Esteban, M. Kowalczyk, and M. Loss · 2014
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Deep pyramidal residual networks
Dongyoon Han, Jiwhan Kim, and Junmo Kim · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Later among the works it cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Later among the works it cites.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
Peter L Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2019
Later among the works it cites.
Higher order concentration of measure
S. G. Bobkov, F. Götze, and H. Sambale · 2019
Later among the works it cites.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Later among the works it cites.
Neural networks learning and memorization with (almost) no over-parameterization, 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding deep learning requires re-thinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
SGD learns over-parameterized networks that provably generalize on linearly separable data
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Cited alongside, same era.
Optimization landscape and expressivity of deep CNNs
Quynh Nguyen and Matthias Hein · 2018
Cited alongside, same era.
Spurious local minima are common in two-layer ReLU neural networks
Itay Safran and Ohad Shamir · 2018
Cited alongside, same era.
A. Daniely · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
S. Du, J. Lee, H. Li, L. Wang, and X. Zhai · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2019
Later among the works it cites.
On connected sublevel sets in deep learning
Q. Nguyen · 2019
Later among the works it cites.
On the loss landscape of a class of deep neural networks with no bad local valleys
Q. Nguyen, M. C. Mukkamala, and M. Hein · 2019
Later among the works it cites.
The effect of network width on stochastic gradient descent and generalization: an empirical study
D. Park, J. Sohl-Dickstein, Q. Le, and S. Smith · 2019
Later among the works it cites.
Small nonlinearities in activation functions create bad local minima in neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2019
Later among the works it cites.
Small ReLU networks are powerful memorizers: a tight analysis of memorization capacity
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2019
Later among the works it cites.
An improved analysis of training over-parameterized deep neural networks
D. Zou and Q. Gu · 2019
Later among the works it cites.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks
Z. Ji and M. Telgarsky · 2020
Closest in time.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Closest in time.