Fetching the paper…
Reading the bibliography…
How can neural networks such as ResNet efficiently learn CIFAR-10 with test accuracy more than 96%, while other methods, especially kernel methods, fall relatively behind? Can we more provide theoretical justifications for this gap? Recently, there is an influential line of work relating neural networks to kernels in the over-parameterized regime, proving they can learn certain concept class that is also learnable by kernels with similar test error.
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Can SGD learn recurrent neural networks with provable generalization?
Zeyuan Allen-Zhu and Yuanzhi Li · 1902
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
Backward Feature Correction: How Deep Learning Performs Deep Learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2001
Earlier work this paper cites.
Feature purification: How adversarial training performs robust deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2005
Earlier work this paper cites.
Non-asymptotic theory of random matrices: extreme singular values
Mark Rudelson and Roman Vershynin · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Recovery guarantee of non-negative matrix factorization via alternating updates
Yuanzhi Li, Yingyu Liang, and Andrej Risteski · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Cited alongside, same era.
Diversity leads to generalization in neural networks
Bo Xie, Yingyu Liang, and Le Song · 2016
Cited alongside, same era.
l1-regularized neural networks are improperly learnable in polynomial time
Yuchen Zhang, Jason D Lee, and Michael I Jordan · 2016
Cited alongside, same era.
Theoretical properties of the global optimizer of two layer neural network
Digvijay Boob and Guanghui Lan · 2017
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Later among the works it cites.
Basic tail and concentration bounds
Martin J. Wainwright · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2017
Cited alongside, same era.
Provable alternating gradient descent for non-negative matrix factorization with strong correlations
Yuanzhi Li and Yingyu Liang · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2017
Cited alongside, same era.
Yuandong Tian · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Cited alongside, same era.
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2018
Later among the works it cites.
Polynomial convergence of gradient descent for training one-hidden-layer neural networks
Santosh Vempala and John Wilmes · 2018
Later among the works it cites.
On the margin theory of feedforward neural networks
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Later among the works it cites.
Learning two-layer neural networks with symmetric inputs
Rong Ge, Rohith Kuditipudi, Zhize Li, and Xiang Wang · 2019
Closest in time.
CS229T/STAT231: Statistical Learning Theory (Fall 2017)
Tengyu Ma · 2019
Closest in time.
Greg Yang · 2019
Closest in time.
When can wasserstein gans minimize wasserstein distance?
Yuanzhi Li and Zehao Dou · 2020
Closest in time.