Fetching the paper…
Reading the bibliography…
Recently, a spate of papers have provided positive theoretical results for training over-parameterized neural networks (where the network size is larger than what is needed to achieve low error).
Bases of tensor products of banach spaces
B. R. Gelbaum, J. G. De Lamadrid, et al · 1961
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
A. R. Barron · 1993
Earlier work this paper cites.
The concentration of measure phenomenon
M. Ledoux · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
Uniform approximation of functions with random bases
A. Rahimi and B. Recht · 2008
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
A. Rahimi and B. Recht · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Learning polynomials with neural networks
A. Andoni, R. Panigrahy, G. Valiant, and L. Zhang · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
R. Livni, S. Shalev-Shwartz, and O. Shamir · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
A. Daniely, R. Frostig, and Y. Singer · 2016
Earlier work this paper cites.
The landscape of empirical risk for non-convex losses
S. Mei, Y. Bai, and A. Montanari · 2016
Cited alongside, same era.
On the quality of the initial basin in overspecified neural networks
I. Safran and O. Shamir · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz · 2017
Cited alongside, same era.
SGD learns the conjugate kernel class of the network
A. Daniely · 2017
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
S. S. Du and J. D. Lee · 2018
Later among the works it cites.
Approximation by combinations of ReLU and squared ReLU ridge functions with l 1 l^{1} and l 0 l^{0} controls
J. M. Klusowski and A. R. Barron · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Later among the works it cites.
Distribution-specific hardness of learning neural networks
O. Shamir · 2018
Later among the works it cites.
Random ReLU features: Universality, approximation, and composition
Y. Sun, A. Gilbert, and A. Tewari · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Li, T. Ma, and H. Zhang · 2017
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
I. Safran and O. Shamir · 2017
Cited alongside, same era.
Learning relus via gradient descent
M. Soltanolkotabi · 2017
Cited alongside, same era.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
D. Soudry and E. Hoffer · 2017
Cited alongside, same era.
Over-parameterization improves generalization in the xor detection problem
A. Brutzkus and A. Globerson · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Z. Allen-Zhu, Y. Li, and Y. Liang
Cited in the paper.
G. Wang, G. B. Giannakis, and J. Chen · 2018
Later among the works it cites.
Can SGD learn recurrent neural networks with provable generalization?
Z. Allen-Zhu and Y. Li · 2019
Closest in time.
A generalization theory of gradient descent for learning over-parameterized deep ReLU networks
Y. Cao and Q. Gu · 2019
Closest in time.
Linearized two-layers neural networks in high dimension
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2019
Closest in time.
Linearized two-layers neural networks in high dimension
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2021
Closest in time.