Fetching the paper…
Reading the bibliography…
Deep networks are often considered to be more expressive than shallow ones in terms of approximation.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Bayesian learning for neural networks
Radford M Neal · 1996
Earlier work this paper cites.
Approximation theory of the mlp model in neural networks
Allan Pinkus · 1999
Earlier work this paper cites.
Regularization with dot-product kernels
Alex J Smola, Zoltan L Ovari, and Robert C Williamson · 2001
Earlier work this paper cites.
On the mathematical foundations of learning
Felipe Cucker and Steve Smale · 2002
Earlier work this paper cites.
Classical and quantum orthogonal polynomials in one variable , volume 13
Mourad Ismail · 2005
Earlier work this paper cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2005
Earlier work this paper cites.
Mercer’s theorem, feature maps, and smoothing
Ha Quang Minh, Partha Niyogi, and Yuan Yao · 2006
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Earlier work this paper cites.
Analytic combinatorics
Philippe Flajolet and Robert Sedgewick · 2009
Earlier work this paper cites.
The spectrum of kernel random matrices
Noureddine El Karoui · 2010
Earlier work this paper cites.
Spherical harmonics and approximations on the unit sphere: an introduction , volume 2044
Kendall Atkinson and Weimin Han · 2012
Earlier work this paper cites.
Sharp analysis of low-rank kernel matrix approximations
Francis Bach · 2013
Earlier work this paper cites.
Sharp estimates for eigenvalues of integral operators generated by dot product kernels on the sphere
Douglas Azevedo and Valdir Antonio Menegatto · 2014
Earlier work this paper cites.
Spherical harmonics in p dimensions
Costas Efthimiou and Christopher Frye · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Cited alongside, same era.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Cited alongside, same era.
Deep vs. shallow networks: An approximation theory perspective
Hrushikesh N Mhaskar and Tomaso Poggio · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Cited alongside, same era.
Depth separation for neural networks
Amit Daniely · 2017
Cited alongside, same era.
Towards understanding the spectral bias of deep learning
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu · 2019
Later among the works it cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Later among the works it cites.
Quadratic suffices for over-parametrization via matrix chernoff bound
Zhao Song and Xin Yang · 2019
Later among the works it cites.
A fine-grained spectral perspective on neural networks
Greg Yang and Hadi Salman · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco · 2017
Cited alongside, same era.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2017
Cited alongside, same era.
Error bounds for approximations with deep relu networks
Dmitry Yarotsky · 2017
Cited alongside, same era.
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Later among the works it cites.
Backward feature correction: How deep learning performs deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2020
Closest in time.
Frequency bias in neural networks for input of non-uniform density
Ronen Basri, Meirav Galun, Amnon Geifman, David Jacobs, Yoni Kasten, and Shira Kritchman · 2020
Closest in time.
Training (overparametrized) neural networks in near-linear time
Jan van den Brand, Binghui Peng, Zhao Song, and Omri Weinstein · 2020
Closest in time.
Towards understanding hierarchical learning: Benefits of neural representations
Minshuo Chen, Yu Bai, Jason D Lee, Tuo Zhao, Huan Wang, Caiming Xiong, and Richard Socher · 2020
Closest in time.
Spectra of the conjugate kernel and neural tangent kernel for linear-width neural networks
Zhou Fan and Zhichao Wang · 2020
Closest in time.
On the similarity between the laplace and neural tangent kernels
Amnon Geifman, Abhay Yadav, Yoni Kasten, Meirav Galun, David Jacobs, and Ronen Basri · 2020
Closest in time.
Generalized leverage score sampling for neural networks
Jason D Lee, Ruoqi Shen, Zhao Song, Mengdi Wang, et al · 2020
Closest in time.
On the risk of minimum-norm interpolants and restricted lower isometry of kernels
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai · 2020
Closest in time.
Risk bounds for multi-layer perceptrons through spectra of integral operators
Meyer Scetbon and Zaid Harchaoui · 2020
Closest in time.
Nonparametric regression using deep neural networks with relu activation function
Johannes Schmidt-Hieber et al · 2020
Closest in time.
Deep neural tangent kernel and laplace kernel have the same rkhs
Lin Chen and Sheng Xu · 2021
Closest in time.