Can SGD Learn Recurrent Neural Networks with Provable Generalization?
Original
Zeyuan Allen-Zhu and Yuanzhi Li · 1902
Earlier work this paper cites.
What Can ResNet Learn Efficiently, Going Beyond Kernels?
Original
Zeyuan Allen-Zhu and Yuanzhi Li · 1905
Earlier work this paper cites.
Matrix perturbation theory
Gilbert W Stewart · 1990
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Smoothed analysis of algorithms: Why the simplex algorithm usually takes polynomial time
Daniel A Spielman and Shang-Hua Teng · 2004
Earlier work this paper cites.
A spectral algorithm for learning mixture models
Santosh Vempala and Grant Wang · 2004
Earlier work this paper cites.
K-means has polynomial smoothed complexity
David Arthur, Bodo Manthey, and Heiko Röglin · 2009
Earlier work this paper cites.
An elementary proof of anti-concentration of polynomials in gaussian variables
Shachar Lovett · 2010
Earlier work this paper cites.
Settling the polynomial learnability of mixtures of gaussians
Ankur Moitra and Gregory Valiant · 2010
Earlier work this paper cites.
A spectral algorithm for latent dirichlet allocation
Anima Anandkumar, Dean P Foster, Daniel J Hsu, Sham M Kakade, and Yi-Kai Liu · 2012
Earlier work this paper cites.
Concentration and moment inequalities for polynomials of independent random variables
Warren Schudy and Maxim Sviridenko · 2012
Earlier work this paper cites.
Tensor decompositions for learning latent variable models
Animashree Anandkumar, Rong Ge, Daniel Hsu, Sham M Kakade, and Matus Telgarsky · 2014
Earlier work this paper cites.
New algorithms for learning incoherent and overcomplete dictionaries
Sanjeev Arora, Rong Ge, and Ankur Moitra · 2014
Earlier work this paper cites.
Smoothed analysis of tensor decompositions
Aditya Bhaskara, Moses Charikar, Ankur Moitra, and Aravindan Vijayaraghavan · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Original
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Earlier work this paper cites.
Lazysvd: Even faster svd decomposition yet without agonizing pain
Zeyuan Allen-Zhu and Yuanzhi Li · 2016
Earlier work this paper cites.
A latent variable model approach to pmi-based word embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2016
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Recovery guarantee of non-negative matrix factorization via alternating updates
Yuanzhi Li, Yingyu Liang, and Andrej Risteski · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Original
Daniel Soudry and Yair Carmon · 2016
Earlier work this paper cites.
Diversity leads to generalization in neural networks
Original
Bo Xie, Yingyu Liang, and Le Song · 2016
Earlier work this paper cites.
Doubly accelerated methods for faster cca and generalized eigendecomposition
Zeyuan Allen-Zhu and Yuanzhi Li · 2017
Earlier work this paper cites.
First efficient convergence for streaming k-pca: a global, gap-free, and near-optimal rate
Zeyuan Allen-Zhu and Yuanzhi Li · 2017
Earlier work this paper cites.
Wasserstein gan
Original
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Earlier work this paper cites.
Generalization and equilibrium in generative adversarial nets (gans)
Sanjeev Arora, Rong Ge, Yingyu Liang, Tengyu Ma, and Yi Zhang · 2017
Earlier work this paper cites.
Theoretical properties of the global optimizer of two layer neural network
Original
Digvijay Boob and Guanghui Lan · 2017
Earlier work this paper cites.