Fetching the paper…
Reading the bibliography…
We study the cost of overfitting in noisy kernel ridge regression (KRR), which we define as the ratio between the test error of the interpolating ridgeless model and the test error of the optimally-tuned model.
Functions of positive and negative type, and their connection the theory of integral equations
James Mercer · 1909
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vladimir Vapnik and Alexey Chervonenkis · 1971
Earlier work this paper cites.
Decision theoretic generalizations of the pac model for neural net and other learning applications
David Haussler · 1992
Earlier work this paper cites.
Some extensions of an inequality of vapnik and chervonenkis
Dmitry Panchenko · 2002
Earlier work this paper cites.
Mercer’s theorem, feature maps, and smoothing
Ha Quang Minh, Partha Niyogi, and Yuan Yao · 2006
Earlier work this paper cites.
Optimistic rates for learning with a smooth loss, 2010
Nathan Srebro, Karthik Sridharan, and Ambuj Tewari · 2010
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Reconciling modern machine learning practice and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler · 2020
Earlier work this paper cites.
On the similarity between the laplace and neural tangent kernels
Amnon Geifman, Abhay Yadav, Yoni Kasten, Meirav Galun, David Jacobs, and Basri Ronen · 2020
Cited alongside, same era.
Kernel alignment risk estimator: Risk prediction from training data
Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler, and Franck Gabriel · 2020
Cited alongside, same era.
Harmless interpolation of noisy data in regression
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai · 2020
Cited alongside, same era.
Benign overfitting in ridge regression
Alexander Tsigler and Peter L. Bartlett · 2020
Cited alongside, same era.
On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression
Denny Wu and Ji Xu · 2020
Cited alongside, same era.
On uniform convergence and low-norm interpolation learning
Lijia Zhou, Danica J. Sutherland, and Nathan Srebro · 2020
James B Simon, Madeline Dickens, Dhruva Karkada, and Michael R. DeWeese · 2021
Later among the works it cites.
Tight bounds for minimum l1-norm interpolation of noisy data
Guillaume Wang, Konstantin Donhauser, and Fanny Yang · 2021
Later among the works it cites.
Optimistic rates: A unifying theory for interpolation learning and regularization in linear regression
Lijia Zhou, Frederic Koehler, Danica J. Sutherland, and Nathan Srebro · 2021
Later among the works it cites.
Dimension free ridge regression
Chen Cheng and Andrea Montanari · 2022
Later among the works it cites.
Fast rates for noisy interpolation require rethinking the effects of inductive bias
Konstantin Donhauser, Nicolo Ruggeri, Stefan Stojanovic, and Fanny Yang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2021
Cited alongside, same era.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Cited alongside, same era.
Uniform convergence of interpolators: Gaussian width, norm bounds and benign overfitting
Frederic Koehler, Lijia Zhou, Danica J. Sutherland, and Nathan Srebro · 2021
Cited alongside, same era.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Bruno Loureiro, Cedric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mezard, and Lenka Zdeborová · 2021
Cited alongside, same era.
A theory of high dimensional regression with arbitrary correlations between input features and target functions: Sample complexity, multiple descent curves and a hierarchy of phase transitions
Gabriel C. Mel and Surya Ganguli · 2021
Cited alongside, same era.
Asymptotics of ridge(less) regression under general source condition
Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco · 2021
Cited alongside, same era.
Benign, tempered, or catastrophic: A taxonomy of overfitting
Neil Rohit Mallinar, James B Simon, Amirhesam Abedsoltan, Parthe Pandit, Mikhail Belkin, and Preetum Nakkiran · 2022
Later among the works it cites.
Generalization error of random feature and kernel methods: Hypercontractivity and kernel matrix concentration
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2022
Later among the works it cites.
Theodor Misiakiewicz · 2022
Later among the works it cites.
More than a toy: Random matrix models predict how real-world neural representations generalize
Alexander Wei, Wei Hu, and Jacob Steinhardt · 2022
Later among the works it cites.
A non-asymptotic moreau envelope theory for high-dimensional generalized linear models
Lijia Zhou, Frederic Koehler, Pragya Sur, Danica J. Sutherland, and Nathan Srebro · 2022
Later among the works it cites.
High-dimensional analysis of double descent for linear regression with random projections
Francis Bach · 2023
Closest in time.