Fetching the paper…
Reading the bibliography…
We prove that a single step of gradient decent over depth two network, with $q$ hidden neurons, starting from orthogonal initialization, can memorize $\Omega\left(\frac{dq}{\log^4(d)}\right)$ independent and randomly labeled Gaussians in $\mathbb{R}^d$.
Samet Oymak and Mahdi Soltanolkotabi · 1902
Earlier work this paper cites.
Learning polynomials with neural networks
A. Andoni, R. Panigrahy, G. Valiant, and L. Zhang · 2014
Earlier work this paper cites.
Probability in high dimension
Ramon van Handel · 2014
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2016
Earlier work this paper cites.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Earlier work this paper cites.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2018
Cited alongside, same era.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Ziwei Ji and Matus Telgarsky · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
Chao Ma, Lei Wu, et al · 2019
Later among the works it cites.
Quadratic suffices for over-parametrization via matrix chernoff bound
Zhao Song and Xin Yang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuan Cao and Quanquan Gu · 2019
Cited alongside, same era.
Neural networks learning and memorization with (almost) no over-parameterization
Amit Daniely · 2019
Cited alongside, same era.
Mildly overparametrized neural nets can memorize training data efficiently
Rong Ge, Runzhe Wang, and Haoyu Zhao · 2019
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang
Cited in the paper.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song
Cited in the paper.
Roman Vershynin · 2019
Later among the works it cites.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Later among the works it cites.