Fetching the paper…
Reading the bibliography…
In this paper, we provide a precise characterization of generalization properties of high dimensional kernel ridge regression across the under- and over-parameterized regimes, depending on whether the number of training data n exceeds the feature dimension d.
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, Dropout: a simple way to prevent neural networks from overfitting , Journal of Machine Learning Research 15
1958
Earlier work this paper cites.
Vladimir A Marčenko and Leonid Andreevich Pastur, Distribution of eigenvalues for some sets of random matrices , Mathematics of the USSR-Sbornik 1
1967
Earlier work this paper cites.
Felix Xinnan Yu, Ananda Theertha Suresh, Krzysztof Choromanski, Daniel Holtmannrice, and Sanjiv Kumar, Orthogonal random features , Advances in Neural Information Processing Systems, 2016, pp. 1975–1983
1983
Earlier work this paper cites.
Yann Lecun, Leon Bottou, Yoshua Bengio, and Patrick Haffner, Gradient-based learning applied to document recognition , the IEEE 86
1998
Earlier work this paper cites.
Felipe Cucker and Steve Smale, On the mathematical foundations of learning , Bulletin of the American mathematical society 39
2002
Earlier work this paper cites.
Johan A.K. Suykens, Tony Van Gestel, Jos De Brabanter, Bart De Moor, and Joos Vandewalle, Least squares support vector machines , World Scientific, 2002
2002
Earlier work this paper cites.
Holger Wendland, Scattered data approximation , vol. 17, Cambridge university press, 2004
2004
Earlier work this paper cites.
Felipe Cucker and Dingxuan Zhou, Learning theory: an approximation theory viewpoint , vol. 24, Cambridge University Press, 2007
2007
Earlier work this paper cites.
Ingo Steinwart and Clint Scovel, Fast rates for support vector machines using Gaussian kernels , Annals of Statistics 35
2007
Earlier work this paper cites.
Steve Smale and Ding-Xuan Zhou, Learning theory estimates via integral operators and their approximations , Constructive Approximation 26
2007
Earlier work this paper cites.
Chih-Chung Chang, LibSVM data: Classification, regression, and multi-label , http://www. csie. ntu. edu. tw/˜cjlin/libsvmtools/datasets/ (2008)
2008
Earlier work this paper cites.
Ingo Steinwart and Christmann Andreas, Support vector machines , Springer Science and Business Media, 2008
2008
Earlier work this paper cites.
Christine De Mol, Ernesto De Vito, and Lorenzo Rosasco, Elastic-net regularization in learning theory , Journal of Complexity 25
2009
Earlier work this paper cites.
Gilles Blanchard and Nicole Krämer, Optimal learning rates for kernel conjugate gradient regression , Advances in Neural Information Processing Systems, 2010, pp. 226–234
2010
Earlier work this paper cites.
Noureddine El Karoui, The spectrum of kernel random matrices , Annals of Statistics 38
2010
Earlier work this paper cites.
Cheng Wang and Ding-Xuan Zhou, Optimal learning rates for least squares regularized regression with unbounded sampling , Journal of Complexity 27
2011
Earlier work this paper cites.
Francis Bach, Sharp analysis of low-rank kernel matrix approximations , Conference on Learning Theory, 2013, pp. 185–209
2013
Earlier work this paper cites.
Yuchen Zhang, John Duchi, and Martin Wainwright, Divide and conquer kernel ridge regression , Conference on Learning Theory, 2013, pp. 592–617
2013
Earlier work this paper cites.
Moritz Hardt, Ben Recht, and Yoram Singer, Train faster, generalize better: Stability of stochastic gradient descent , International Conference on Machine Learning, 2016, pp. 1225–1234
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Haim Avron, Michael Kapralov, Cameron Musco, Christopher Musco, Ameya Velingker, and Amir Zandieh, Random Fourier features for kernel ridge regression: Approximation bounds and statistical guarantees , the 34th International Conference on Machine Learning, 2017, pp. 253–262
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
Zheng-Chu Guo, Lei Shi, and Qiang Wu, Learning theory of distributed regression with bias corrected regularization kernel network , Journal of Machine Learning Research 18
2017
Cited alongside, same era.
Shao-Bo Lin, Xin Guo, and Ding-Xuan Zhou, Distributed learning with regularized least squares , Journal of Machine Learning Research 18
2017
Cited alongside, same era.
Alessandro Rudi and Lorenzo Rosasco, Generalization properties of learning with random features , Advances in Neural Information Processing Systems, 2017, pp. 3215–3225
2017
Cited alongside, same era.
Edgar Dobriban and Stefan Wager, High-dimensional asymptotics of prediction: Ridge regression and classification , Annals of Statistics 46
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
Khalil Elkhalil, Abla Kammoun, Xiangliang Zhang, Mohamed-Slim Alouini, and Tareq Al-Naffouri, Risk convergence of centered kernel ridge regression with large dimensional data , IEEE Transactions on Signal Processing 68
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Zhenyu Liao and Romain Couillet, On the spectrum of random features maps of high dimensional data , the International Conference on Machine Learning, 2018, pp. 3063–3071
2018
Cited alongside, same era.
Alnur Ali, J Zico Kolter, and Ryan J Tibshirani, A continuous-time view of early stopping for least squares regression , International Conference on Artificial Intelligence and Statistics, 2019, pp. 1370–1378
2019
Cited alongside, same era.
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off , the National Academy of Sciences 116
2019
Cited alongside, same era.
Mikhail Belkin, Alexander Rakhlin, and Alexandre B Tsybakov, Does data interpolation contradict statistical optimality? , International Conference on Artificial Intelligence and Statistics, 2019, pp. 1611–1619
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Linearized two-layers neural networks in high dimension , Annals of Statistics (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová, Generalisation error in learning with random features and the hidden manifold model , International Conference on Machine Learning, 2020, pp. 3452–3462
2020
Closest in time.
Arthur Jacot, Berfin Şimşek, Francesco Spadaro, Clément Hongler, and Franck Gabriel, Implicit regularization of random feature models , International Conference on Machine Learning, 2020, pp. 4631–4640
2020
Closest in time.
Dmitry Kobak, Jonathan Lomond, and Benoit Sanchez, The optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization , Journal of Machine Learning Research 21
2020
Closest in time.
Zhenyu Liao, Romain Couillet, and Michael Mahoney, A random matrix analysis of random fourier features: beyond the gaussian kernel, a precise phase transition, and the corresponding double descent , Neural Information Processing Systems, 2020
2020
Closest in time.
Sifan Liu and Edgar Dobriban, Ridge regression: Structure, cross-validation, and sketching , International Conference on Learning Representations, 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
Daniel LeJeune, Hamid Javadi, and Richard Baraniuk, The implicit regularization of ordinary least squares ensembles , International Conference on Artificial Intelligence and Statistics, 2020, pp. 3525–3535
2020
Closest in time.
Tengyuan Liang and Alexander Rakhlin, Just interpolate: Kernel “ridgeless” regression can generalize , Annals of Statistics 48
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
Denny Wu and Ji Xu, On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression , Advances in Neural Information Processing Systems, 2020, pp. 1–11
2020
Closest in time.
Zitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt, and Yi Ma, Rethinking bias-variance trade-off for generalization of neural networks , the International Conference on Machine Learning, 2020
2020
Closest in time.
Fanghui Liu, Lei Shi, Xiaolin Huang, Jie Yang, and Johan A.K. Suykens, Analysis of regularized least squares in reproducing kernel kreĭn spaces , Machine Learning (2021), 1–20
2021
Closest in time.
Alessandro Rudi, Guillermo D Canas, and Lorenzo Rosasco, On the sample complexity of subspace learning , Advances in Neural Information Processing Systems, 2013, pp. 2067–2075
2075
Closest in time.