Fetching the paper…
Reading the bibliography…
We study the spectrum of inner-product kernel matrices, i.e., $n \times n$ matrices with entries $h (\langle \textbf{x}_i ,\textbf{x}_j \rangle/d)$ where the $( \textbf{x}_i)_{i \leq n}$ are i.i.d.~random covariates in $\mathbb{R}^d$.
Szegö, Gabor, Orthogonal polynomials , vol. 23, American Mathematical Soc., 1939
1939
Earlier work this paper cites.
Vladimir Alexandrovich Marchenko and Leonid Andreevich Pastur, Distribution of eigenvalues for some sets of random matrices , Matematicheskii Sbornik 114
1967
Earlier work this paper cites.
Aline Bonami, Etude des coefficients de Fourier des fonctions de L p ( G ) L^{p}(G) , Annales de l’institut Fourier, vol. 20, 1970, pp. 335–402
1970
Earlier work this paper cites.
William Beckner, Inequalities in Fourier analysis , Annals of Mathematics (1975), 159–182
1975
Earlier work this paper cites.
Leonard Gross, Logarithmic sobolev inequalities , American Journal of Mathematics 97
1975
Earlier work this paper cites.
Andrea Caponnetto and Ernesto De Vito, Optimal rates for the regularized least-squares algorithm , Foundations of Computational Mathematics 7
2007
Earlier work this paper cites.
Zhidong Bai and Wang Zhou, Large sample covariance matrices without independence structures in columns , Statistica Sinica (2008), 425–442
2008
Earlier work this paper cites.
Noureddine El Karoui, The spectrum of kernel random matrices , The Annals of Statistics 38
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
Radoslaw Adamczak, On the marchenko-pastur and circular laws for some classes of random matrices with dependent entries , Electronic Journal of Probability 16
2011
Earlier work this paper cites.
Alain Berlinet and Christine Thomas-Agnan, Reproducing kernel hilbert spaces in probability and statistics , Springer Science & Business Media, 2011
2011
Earlier work this paper cites.
Theodore S Chihara, An introduction to orthogonal polynomials , Courier Corporation, 2011
2011
Earlier work this paper cites.
Leonid Andreevich Pastur and Mariya Shcherbina, Eigenvalue distribution of large random matrices , no. 171, American Mathematical Soc., 2011
2011
Earlier work this paper cites.
John S Avery, Hyperspherical harmonics: applications in quantum theory , vol. 5, Springer Science & Business Media, 2012
2012
Earlier work this paper cites.
Sean O’Rourke, A note on the marchenko-pastur law for a class of random matrices with dependent entries , Electronic Communications in Probability 17
2012
Earlier work this paper cites.
Xiuyuan Cheng and Amit Singer, The spectrum of random inner-product kernel matrices , Random Matrices: Theory and Applications 2
2013
Earlier work this paper cites.
Yen Do and Van Vu, The spectrum of random kernel matrices: universality results for rough and varying kernels , Random Matrices: Theory and Applications 2
2013
Earlier work this paper cites.
Feng Dai and Yuan Xu, Spherical harmonics , Approximation theory and harmonic analysis on spheres and balls, Springer, 2013, pp. 1–27
2013
Earlier work this paper cites.
Bloemendal Alex, Laszlo Erdös, Antti Knowles, Horng-Tzer Yau, and Jun Yin, Isotropic local laws for sample covariance and generalized wigner matrices , Electronic Journal of Probability 19
2014
Earlier work this paper cites.
Costas Efthimiou and Christopher Frye, Spherical harmonics in p dimensions , World Scientific, 2014
2014
Cited alongside, same era.
Ryan O’Donnell, Analysis of boolean functions , Cambridge University Press, 2014
2014
Cited alongside, same era.
Pavel Yaskov, Necessary and sufficient conditions for the marchenko-pastur theorem , Electronic Communications in Probability 21
2016
Cited alongside, same era.
Mikhail Belkin, Siyuan Ma, and Soumik Mandal, To understand deep learning we need to understand kernel learning , International Conference on Machine Learning, PMLR, 2018, pp. 541–549
2018
Cited alongside, same era.
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh, Gradient descent provably optimizes over-parameterized neural networks , International Conference on Learning Representations, 2018
2018
Cited alongside, same era.
Dmitry Kobak, Jonathan Lomond, and Benoit Sanchez, The optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization. , J. Mach. Learn. Res. 21
2020
Later among the works it cites.
Zhenyu Liao, Romain Couillet, and Michael W Mahoney, A random matrix analysis of random fourier features: beyond the gaussian kernel, a precise phase transition, and the corresponding double descent , Advances in Neural Information Processing Systems 33
2020
Later among the works it cites.
Tengyuan Liang and Alexander Rakhlin, Just interpolate: Kernel “ridgeless” regression can generalize , The Annals of Statistics 48
2020
Later among the works it cites.
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai, On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels , Conference on Learning Theory, PMLR, 2020, pp. 2683–2711
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arthur Jacot, Franck Gabriel, and Clément Hongler, Neural tangent kernel: Convergence and generalization in neural networks , Advances in neural information processing systems, 2018, pp. 8571–8580
2018
Cited alongside, same era.
Yuanzhi Li and Yingyu Liang, Learning overparameterized neural networks via stochastic gradient descent on structured data , Advances in Neural Information Processing Systems, 2018, pp. 8157–8166
2018
Cited alongside, same era.
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song, On the convergence rate of training recurrent neural networks , Advances in Neural Information Processing Systems 32
2019
Cited alongside, same era.
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off , Proceedings of the National Academy of Sciences 116
2019
Cited alongside, same era.
Lenaic Chizat, Edouard Oyallon, and Francis Bach, On lazy training in differentiable programming , NeurIPS 2019-33rd Conference on Neural Information Processing Systems, 2019, pp. 2937–2947
2019
Cited alongside, same era.
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington, Wide neural networks of any depth evolve as linear models under gradient descent , Advances in neural information processing systems 32
2019
Cited alongside, same era.
Alexander Rakhlin and Xiyu Zhai, Consistency of interpolation with laplace kernels is a high-dimensional phenomenon , Conference on Learning Theory, PMLR, 2019, pp. 2595–2623
2019
Cited alongside, same era.
2020
Later among the works it cites.
Denny Wu and Ji Xu, On the optimal weighted e l l _ 2 ell\_2 regularization in overparameterized linear regression , Advances in Neural Information Processing Systems 33
2020
Later among the works it cites.
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin, Deep learning: a statistical viewpoint , Acta numerica 30
2021
Later among the works it cites.
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan, Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks , Nature communications 12
2021
Later among the works it cites.
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová, Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime , Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Lin Chen, Yifei Min, Mikhail Belkin, and Amin Karbasi, Multiple descent: Design your own generalization curve , Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Fanghui Liu, Zhenyu Liao, and Johan Suykens, Kernel regression in high dimensions: Refined analysis beyond double descent , International Conference on Artificial Intelligence and Statistics, PMLR, 2021, pp. 649–657
2021
Later among the works it cites.
2021
Later among the works it cites.
Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Generalization error of random feature and kernel methods: Hypercontractivity and kernel matrix concentration , Applied and Computational Harmonic Analysis (2021)
2021
Later among the works it cites.
Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Learning with invariances in random features and kernel models , Conference on Learning Theory, PMLR, 2021, pp. 3351–3418
2021
Later among the works it cites.
Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco, Asymptotics of ridge (less) regression under general source condition , International Conference on Artificial Intelligence and Statistics, PMLR, 2021, pp. 3889–3897
2021
Later among the works it cites.
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani, Surprises in high-dimensional ridgeless least squares interpolation , The Annals of Statistics 50
2022
Closest in time.
Song Mei and Andrea Montanari, The generalization error of random features regression: Precise asymptotics and the double descent curve , Communications on Pure and Applied Mathematics 75
2022
Closest in time.