Fetching the paper…
Reading the bibliography…
We consider the problem of learning an unknown function $f_{\star}$ on the $d$-dimensional sphere with respect to the square loss, given i.i.d.
Szegő, Gabor, Orthogonal polynomials , vol. 23, American Mathematical Soc., 1939
1939
Earlier work this paper cites.
George Cybenko, Approximation by superpositions of a sigmoidal function , Mathematics of control, signals and systems 2
1989
Earlier work this paper cites.
Ronald A DeVore, Ralph Howard, and Charles Micchelli, Optimal nonlinear approximation , Manuscripta mathematica 63
1989
Earlier work this paper cites.
David L Donoho and Iain M Johnstone, Projection-based approximation and a duality with kernel methods , The Annals of Statistics (1989), 58–106
1989
Earlier work this paper cites.
Kurt Hornik, Approximation capabilities of multilayer feedforward networks , Neural networks 4
1991
Earlier work this paper cites.
Andrew R Barron, Universal approximation bounds for superpositions of a sigmoidal function , IEEE Transactions on Information theory 39
1993
Earlier work this paper cites.
Hrushikesh Narhar Mhaskar and Charles A Micchelli, Dimension-independent bounds on the degree of approximation by neural networks , IBM Journal of Research and Development 38
1994
Earlier work this paper cites.
Federico Girosi, Michael Jones, and Tomaso Poggio, Regularization theory and neural networks architectures , Neural computation 7
1995
Earlier work this paper cites.
Hrushikesh N Mhaskar, Neural networks for optimal approximation of smooth and analytic functions , Neural computation 8
1996
Earlier work this paper cites.
Radford M Neal, Priors for infinite networks , Bayesian Learning for Neural Networks, Springer, 1996, pp. 29–53
1996
Earlier work this paper cites.
Pencho P Petrushev, Approximation by ridge functions and neural networks , SIAM Journal on Mathematical Analysis 30
1998
Earlier work this paper cites.
VE Maiorov, On best approximation by ridge functions , Journal of Approximation Theory 99
1999
Earlier work this paper cites.
Allan Pinkus, Approximation theory of the mlp model in neural networks , Acta numerica 8
1999
Earlier work this paper cites.
Nello Cristianini, John Shawe-Taylor, et al., An introduction to support vector machines and other kernel-based learning methods , Cambridge University Press, 2000
2000
Earlier work this paper cites.
László Györfi, Michael Kohler, Adam Krzyzak, and Harro Walk, A distribution-free theory of nonparametric regression , Springer Science & Business Media, 2006
2006
Earlier work this paper cites.
Andrea Caponnetto and Ernesto De Vito, Optimal rates for the regularized least-squares algorithm , Foundations of Computational Mathematics 7
2007
Earlier work this paper cites.
Ali Rahimi and Benjamin Recht, Random features for large-scale kernel machines , Advances in neural information processing systems, 2008, pp. 1177–1184
2008
Earlier work this paper cites.
Alexandre B Tsybakov, Introduction to nonparametric estimation , Springer Science & Business Media, 2008
2008
Cited alongside, same era.
Martin Anthony and Peter L Bartlett, Neural network learning: Theoretical foundations , cambridge university press, 2009
2009
Cited alongside, same era.
Noureddine El Karoui, On information plus noise kernel random matrices , The Annals of Statistics 38
2010
Cited alongside, same era.
Alain Berlinet and Christine Thomas-Agnan, Reproducing kernel hilbert spaces in probability and statistics , Springer Science & Business Media, 2011
2011
Cited alongside, same era.
Theodore S Chihara, An introduction to orthogonal polynomials , Courier Corporation, 2011
2011
Cited alongside, same era.
Song Mei, Andrea Montanari, and Phan-Minh Nguyen, A mean field view of the landscape of two-layer neural networks , Proceedings of the National Academy of Sciences (2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro, The implicit bias of gradient descent on separable data , The Journal of Machine Learning Research 19
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Francis Bach, Sharp analysis of low-rank kernel matrix approximations , Conference on Learning Theory, 2013, pp. 185–209
2013
Cited alongside, same era.
Costas Efthimiou and Christopher Frye, Spherical harmonics in p dimensions , World Scientific, 2014
2014
Cited alongside, same era.
Ahmed El Alaoui and Michael W Mahoney, Fast randomized kernel ridge regression with statistical guarantees , Advances in Neural Information Processing Systems, 2015, pp. 775–783
2015
Cited alongside, same era.
2016
Cited alongside, same era.
Alessandro Rudi and Lorenzo Rosasco, Generalization properties of learning with random features , Advances in Neural Information Processing Systems, 2017, pp. 3215–3225
2017
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Later among the works it cites.
2018
Later among the works it cites.
2019
Closest in time.
2019
Closest in time.
Lenaic Chizat, Edouard Oyallon, and Francis Bach, On lazy training in differentiable programming , Advances in Neural Information Processing Systems, 2019, pp. 2933–2943
2019
Closest in time.
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Limitations of lazy training of two-layers neural network , Advances in Neural Information Processing Systems, 2019, pp. 9108–9118
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.