Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Original
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Can SGD Learn Recurrent Neural Networks with Provable Generalization?
Original
Zeyuan Allen-Zhu and Yuanzhi Li · 1902
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Original
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
What Can ResNet Learn Efficiently, Going Beyond Kernels?
Original
Zeyuan Allen-Zhu and Yuanzhi Li · 1905
Earlier work this paper cites.
What Can ResNet Learn Efficiently, Going Beyond Kernels?
Original
Zeyuan Allen-Zhu and Yuanzhi Li · 1905
Earlier work this paper cites.
A theory of the learnable
Leslie Valiant · 1984
Earlier work this paper cites.
Smoothed analysis of algorithms: Why the simplex algorithm usually takes polynomial time
Daniel Spielman and Shang-Hua Teng · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Simple, efficient, and neural algorithms for sparse coding
Sanjeev Arora, Rong Ge, Tengyu Ma, and Ankur Moitra · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Complexity theoretic limitations on learning halfspaces
Amit Daniely · 2016
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Recovery guarantee of non-negative matrix factorization via alternating updates
Yuanzhi Li, Yingyu Liang, and Andrej Risteski · 2016
Earlier work this paper cites.