Fetching the paper…
Reading the bibliography…
This paper establishes rates of universal approximation for the shallow neural tangent kernel (NTK): network weights are only allowed microscopic changes from random initialization, which entails that activations are mostly unchanged, and the network is nearly equivalent to its linearization.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
Yuan Cao and Quanquan Gu · 1902
Earlier work this paper cites.
Samet Oymak and Mahdi Soltanolkotabi · 1902
Earlier work this paper cites.
What can resnet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 1905
Earlier work this paper cites.
A function space view of bounded norm infinite width relu nets: The multivariate case
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro · 1910
Earlier work this paper cites.
Remarques sur un résultat non publié de b. maurey
Gilles Pisier · 1980
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
On the approximate realization of continuous mappings by neural networks
Ken-ichi Funahashi · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Approximation by superposition of sigmoidal and radial basis functions
Hrushikesh N Mhaskar and Charles A Micchelli · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R. Barron · 1993
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Moshe Leshno, Vladimir Ya. Lin, Allan Pinkus, and Shimon Schocken · 1993
Cited alongside, same era.
Approximation and learning of convex superpositions
Leonid Gurvits and Pascal Koiran · 1995
Cited alongside, same era.
Real analysis: modern techniques and their applications
Gerald B. Folland · 1999
Cited alongside, same era.
Scattered Data Approximation
Holger Wendland · 2004
Cited alongside, same era.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Cited alongside, same era.
Kernel methods for deep learning
Youngmin Cho and Lawrence K. Saul · 2009
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Simon S. Du, Jason D. Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2018
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
UC Berkeley Statistics 210B, Lecture Notes: Basic tail and concentration bounds, Jan 2015
Martin J. Wainwright · 2015
Cited alongside, same era.
Error bounds for approximations with deep relu networks
Dmitry Yarotsky · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two Layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach
Cited in the paper.
Later among the works it cites.
A Mean Field View of the Landscape of Two-Layers Neural Networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
On the approximation properties of random relu features
Yitong Sun, Anna Gilbert, and Ambuj Tewari · 2018
Later among the works it cites.
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Closest in time.
The convergence rate of neural networks for learned functions of different frequencies
Ronen Basri, David Jacobs, Yoni Kasten, and Shira Kritchman · 2019
Closest in time.
On the inductive bias of neural tangent kernels
Alberto Bietti and Julien Mairal · 2019
Closest in time.
A Note on Lazy Training in Supervised Differentiable Programming
Lenaic Chizat and Francis Bach · 2019
Closest in time.