Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Original
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Decoupled greedy learning of cnns
Original
Eugene Belilovsky, Michael Eickenberg, and Edouard Oyallon · 1901
Earlier work this paper cites.
Can SGD Learn Recurrent Neural Networks with Provable Generalization?
Original
Zeyuan Allen-Zhu and Yuanzhi Li · 1902
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Original
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
What Can ResNet Learn Efficiently, Going Beyond Kernels?
Original
Zeyuan Allen-Zhu and Yuanzhi Li · 1905
Earlier work this paper cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Original
Yuanzhi Li, Colin Wei, and Tengyu Ma · 1907
Earlier work this paper cites.
Enhanced convolutional neural tangent kernels
Original
Zhiyuan Li, Ruosong Wang, Dingli Yu, Simon S Du, Wei Hu, Ruslan Salakhutdinov, and Sanjeev Arora · 1911
Earlier work this paper cites.
Understanding the difficulty of training transformers
Original
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han · 2004
Earlier work this paper cites.
Feature purification: How adversarial training performs robust deep learning
Original
Zeyuan Allen-Zhu and Yuanzhi Li · 2005
Earlier work this paper cites.
New results for learning noisy parities and halfspaces
Vitaly Feldman, Parikshit Gopalan, Subhash Khot, and Ashok Kumar Ponnuswami · 2006
Earlier work this paper cites.
Very deep transformers for neural machine translation
Original
Xiaodong Liu, Kevin Duh, Liyuan Liu, and Jianfeng Gao · 2008
Earlier work this paper cites.
Learning deep architectures for AI
Yoshua Bengio · 2009
Earlier work this paper cites.
Hierarchical learning: Theory with applications in speech and vision
Jacob V Bouvrie · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
An elementary proof of anti-concentration of polynomials in gaussian variables
Shachar Lovett · 2010
Earlier work this paper cites.
Concentration and moment inequalities for polynomials of independent random variables
Warren Schudy and Maxim Sviridenko · 2012
Earlier work this paper cites.
Provable bounds for learning some deep representations
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Simple, efficient, and neural algorithms for sparse coding
Sanjeev Arora, Rong Ge, Tengyu Ma, and Ankur Moitra · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Original
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
LazySVD: even faster SVD decomposition yet without agonizing pain
Zeyuan Allen-Zhu and Yuanzhi Li · 2016
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.