Fetching the paper…
Reading the bibliography…
We consider the dynamic of gradient descent for learning a two-layer neural network.
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Can SGD Learn Recurrent Neural Networks with Provable Generalization?
Zeyuan Allen-Zhu and Yuanzhi Li · 1902
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
Yuanzhi Li, Colin Wei, and Tengyu Ma · 1907
Earlier work this paper cites.
Enhanced convolutional neural tangent kernels
Zhiyuan Li, Ruosong Wang, Dingli Yu, Simon S Du, Wei Hu, Ruslan Salakhutdinov, and Sanjeev Arora · 1911
Earlier work this paper cites.
Backward feature correction: How deep learning performs deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2001
Earlier work this paper cites.
Feature purification: How adversarial training performs robust deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2005
Earlier work this paper cites.
Analysis of boolean functions
Ryan O’Donnell · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Escaping from saddle points: online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
When are nonconvex problems not scary?
Ju Sun, Qing Qu, and John Wright · 2015
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Earlier work this paper cites.
Diversity leads to generalization in neural networks
Bo Xie, Yingyu Liang, and Le Song · 2016
Earlier work this paper cites.
Analyzing tensor power method dynamics in overcomplete regime
Animashree Anandkumar, Rong Ge, and Majid Janzamin · 2017
Earlier work this paper cites.
Theoretical properties of the global optimizer of two layer neural network
Digvijay Boob and Guanghui Lan · 2017
Earlier work this paper cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Earlier work this paper cites.
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
Simon S Du, Jason D Lee, Yuandong Tian, Barnabas Poczos, and Aarti Singh · 2017
Cited alongside, same era.
On the optimization landscape of tensor decompositions
Rong Ge and Tengyu Ma · 2017
Cited alongside, same era.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017
Cited alongside, same era.
First-order methods almost always avoid saddle points
Jason D Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2017
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Later among the works it cites.
Polynomial convergence of gradient descent for training one-hidden-layer neural networks
Santosh Vempala and John Wilmes · 2018
Later among the works it cites.
On the margin theory of feedforward neural networks
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2018
Later among the works it cites.
Learning one-hidden-layer relu networks via gradient descent
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Provable alternating gradient descent for non-negative matrix factorization with strong correlations
Yuanzhi Li and Yingyu Liang · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2017
Cited alongside, same era.
Yuandong Tian · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Cited alongside, same era.
Learning two layer rectified neural networks in polynomial time
Ainesh Bakshi, Rajesh Jayaram, and David P Woodruff · 2018
Cited alongside, same era.
Xiao Zhang, Yaodong Yu, Lingxiao Wang, and Quanquan Gu · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Later among the works it cites.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Yu Bai and Jason D Lee · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Later among the works it cites.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2019
Later among the works it cites.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
Optimization for deep learning: theory and algorithms
Ruoyu Sun · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Later among the works it cites.
Greg Yang · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance, 2020
Jeff Z. HaoChen, Colin Wei, Jason D. Lee, and Tengyu Ma · 2020
Closest in time.
When can wasserstein gans minimize wasserstein distance?
Yuanzhi Li and Zehao Dou · 2020
Closest in time.