Fetching the paper…
Reading the bibliography…
Recent theoretical works on over-parameterized neural nets have focused on two aspects: optimization and generalization.
Training a 3-node neural network is np-complete
Avrim Blum and Ronald L Rivest · 1989
Earlier work this paper cites.
On the infeasibility of training neural networks with small squared errors
Van H Vu · 1998
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C D. Freeman and J. Bruna · 2016
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
P. L. Bartlett, D. J. Foster, and M. J. Telgarsky · 2017
Earlier work this paper cites.
Globally optimal gradient descent for a convnet with gaussian inputs
A. Brutzkus and A. Globerson · 2017
Earlier work this paper cites.
The multilinear structure of relu networks
Thomas Laurent and James von Brecht · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Y. Li and Y. Yuan · 2017
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix recovery
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2017
Earlier work this paper cites.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Earlier work this paper cites.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
K. Zhong, Z. Song, P. Jain, P. L Bartlett, and I. S Dhillon · 2017
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz · 2018
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
S. S Du and J. D Lee · 2018
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Later among the works it cites.
Generalization error bounds of gradient descent for learning overparameterized deep ReLU networks
Yuan Cao and Quanquan Gu · 2019
Later among the works it cites.
How much over-parameterization is sufficient to learn deep relu networks?
Zixiang Chen, Yuan Cao, Difan Zou, and Quanquan Gu · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2018
Cited alongside, same era.
Adding one neuron can eliminate all bad local minima
S. Liang, R. Sun, J. D Lee, and R Srikant · 2018
Cited alongside, same era.
Understanding the loss surface of neural networks for binary classification
Shiyu Liang, Ruoyu Sun, Yixuan Li, and Rayadurgam Srikant · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layers neural networks
S. Mei, A. Montanari, and P. Nguyen · 2018
Cited alongside, same era.
On the connection between learning two-layers neural networks and tensor decomposition
Marco Mondelli and Andrea Montanari · 2018
Cited alongside, same era.
Towards understanding the role of over-parametrization in generalization of neural networks
B. Neyshabur, Z. Li, S. Bhojanapalli, Y. LeCun, and N. Srebro · 2018
Cited alongside, same era.
On the loss landscape of a class of deep neural networks with no bad local valleys
Q. Nguyen, M. C. Mukkamala, and M. Hein · 2018
Cited alongside, same era.
Ziwei Ji and Matus Telgarsky · 2019
Later among the works it cites.
Revisiting landscape analysis in deep neural networks: Eliminating decreasing paths to infinity
Shiyu Liang, Ruoyu Sun, and R Srikant · 2019
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Later among the works it cites.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Later among the works it cites.
A law of robustness for two-layers neural networks
Sébastien Bubeck, Yuanzhi Li, and Dheeraj Nagaraj · 2020
Later among the works it cites.
Superpolynomial lower bounds for learning one-layer neural networks using gradient descent
Surbhi Goel, Aravind Gollakota, Zhihan Jin, Sushrut Karmalkar, and Adam Klivans · 2020
Later among the works it cites.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Later among the works it cites.
The implicit and explicit regularization effects of dropout
Colin Wei, Sham Kakade, and Tengyu Ma · 2020
Later among the works it cites.