Fetching the paper…
Reading the bibliography…
The success of over-parameterized neural networks trained to near-zero training error has caused great interest in the phenomenon of benign overfitting, where estimators are statistically consistent even though they interpolate noisy training data.
Über die Abgrenzung der Eigenwerte einer Matrix
Semjon Gerschgorin · 1931
Earlier work this paper cites.
Zur Abschätzung von Matrizennormen
Hans Richter · 1958
Earlier work this paper cites.
On the trace of matrix products
Leon Mirsky · 1959
Earlier work this paper cites.
Eigenvalues and s-Numbers
Albrecht Pietsch · 1987
Earlier work this paper cites.
Entropy, Compactness and the Approximation of Operators
Bernd Carl and Irmtraud Stephani · 1990
Earlier work this paper cites.
Besov spaces on domains in ℝ d \mathbb{R}^{d}
Ronald A DeVore and Robert C Sharpley · 1993
Earlier work this paper cites.
Function Spaces, Entropy Numbers, Differential Operators
David E. Edmunds and Hans Triebel · 1996
Earlier work this paper cites.
Priors for Infinite Networks
Radford M. Neal · 1996
Earlier work this paper cites.
The Hilbert kernel regression estimate
Luc Devroye, Laszlo Györfi, and Adam Krzyżak · 1998
Earlier work this paper cites.
An Introduction to Banach Space Theory
Robert E. Megginson · 1998
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Béatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Consistency of support vector machines and other regularized kernel machines
Ingo Steinwart · 2001
Earlier work this paper cites.
Sobolev Spaces
Robert A. Adams and John J.F. Fournier · 2003
Earlier work this paper cites.
Gaussian Processes for Machine Learning
Carl Edward Rasmussen and Christopher K. I. Williams · 2005
Earlier work this paper cites.
Scattered Data Approximation
Holger Wendland · 2005
Earlier work this paper cites.
Concentration inequalities and martingale inequalities: a survey
Fan Chung and Linyuan Lu · 2006
Earlier work this paper cites.
On early stopping in gradient descent learning
Yuan Yao, Lorenzo Rosasco, and Andrea Caponnetto · 2007
Earlier work this paper cites.
Support Vector Machines
Ingo Steinwart and Andreas Christmann · 2008
Earlier work this paper cites.
Optimal rates for regularized least squares regression
Ingo Steinwart, Don R. Hush, and Clint Scovel · 2009
Earlier work this paper cites.
Bounding standard gaussian tail probabilities
Lutz Duembgen · 2010
Earlier work this paper cites.
Essentials of integration theory for analysis
Daniel W. Stroock et al · 2011
Earlier work this paper cites.
Hitchhiker’s guide to the fractional Sobolev spaces
Eleonora Di Nezza, Giampiero Palatucci, and Enrico Valdinoci · 2012
Earlier work this paper cites.
Non-homogeneous boundary value problems and applications: Vol. 1
Jacques Louis Lions and Enrico Magenes · 2012
Earlier work this paper cites.
Mercer’s theorem on general domains: on the interaction between measures, kernels, and RKHSs
Ingo Steinwart and Clint Scovel · 2012
Earlier work this paper cites.
Strictly and non-strictly positive definite functions on spheres
Tilmann Gneiting · 2013
Earlier work this paper cites.
Matrix Analysis
Roger A. Horn and Charles R. Johnson · 2013
Earlier work this paper cites.
Sobolev spaces on Riemannian manifolds with bounded geometry: General coordinates traces
Cornelia Schneider and Nadine Große · 2013
Earlier work this paper cites.
The singular value decomposition of compact operators on Hilbert spaces, 2014
Jordan Bell · 2014
Earlier work this paper cites.
Early stopping and non-parametric regression: An optimal data-dependent stopping rule
Garvesh Raskutti, Martin J. Wainwright, and Bin Yu · 2014
Earlier work this paper cites.
Sobolev spaces, their generalizations and elliptic problems in smooth and Lipschitz domains
Mikhail S Agranovich · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Spherical Radial Basis Functions, Theory and Applications
Simon Hubbert, Quoc Le Gia, and Tanya Morton · 2015
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
A short note on the comparison of interpolation widths, entropy numbers, and Kolmogorov widths
Ingo Steinwart · 2017
Cited alongside, same era.
Early stopping for kernel boosting algorithms: A general analysis with localized complexities
Yuting Wei, Fanny Yang, and Martin J Wainwright · 2017
Cited alongside, same era.
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Cited alongside, same era.
Neural Tangent Kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Gaussian processes and kernel methods: A review on connections and equivalences
Motonobu Kanagawa, Philipp Hennig, Dino Sejdinovic, and Bharath K. Sriperumbudur · 2018
Cited alongside, same era.
Reproducing kernel hilbert spaces on manifolds: Sobolev and diffusion spaces
Ernesto De Vito, Nicole Mücke, and Lorenzo Rosasco · 2021
Later among the works it cites.
How rotational invariance of common kernels prevents generalization in high dimensions
Konstantin Donhauser, Mingqi Wu, and Fanny Yang · 2021
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Later among the works it cites.
On the universality of the double descent peak in ridgeless regression
David Holzmüller · 2021
Later among the works it cites.
Regularization matters: A nonparametric perspective on overparametrized neural network
Tianyang Hu, Wenjia Wang, Cong Lin, and Guang Cheng · 2021
Later among the works it cites.
Early-stopped neural networks are consistent
Ziwei Ji, Justin D. Li, and Matus Telgarsky · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep neural networks as Gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Alexander G. de G. Matthews, Jiri Hron, Mark Rowland, Richard E. Turner, and Zoubin Ghahramani · 2018
Cited alongside, same era.
Superconvergence of kernel-based interpolation
Robert Schaback · 2018
Cited alongside, same era.
Relative concentration bounds for the spectrum of kernel matrices
Ernesto Araya Valdivia · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Cited alongside, same era.
Generalized Inverses: Theory and Computations
Guorong Wang, Yimin Wei, Sanzheng Qiao, Peng Lin, and Yuzhuo Chen · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, Russ R. Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Later among the works it cites.
The future is log-gaussian: Resnets and their infinite-depth-and-width limit at initialization
Mufan Li, Mihai Nica, and Dan Roy · 2021
Later among the works it cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu, and Anant Sahai · 2021
Later among the works it cites.
Uniform generalization bounds for overparameterized neural networks
Sattar Vakili, Michael Bromberg, Jezabel Garcia, Da-shan Shiu, and Alberto Bernacchia · 2021
Later among the works it cites.
Tensor programs iv: Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
Kernel interpolation in sobolev spaces is not consistent in low dimensions
Simon Buchholz · 2022
Later among the works it cites.
Fast rates for noisy interpolation require rethinking the effect of inductive bias
Konstantin Donhauser, Nicolò Ruggeri, Stefan Stojanovic, and Fanny Yang · 2022
Later among the works it cites.
Benign overfitting without linearity: Neural network classifiers trained by gradient descent for noisy linear data
Spencer Frei, Niladri S Chatterji, and Peter Bartlett · 2022
Later among the works it cites.
A universal trade-off between the model size, test loss, and training loss of linear predictors
Nikhil Ghosh and Mikhail Belkin · 2022
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J. Tibshirani · 2022
Later among the works it cites.
Training two-layer relu networks with gradient descent is inconsistent
David Holzmüller and Ingo Steinwart · 2022
Later among the works it cites.
Benign, tempered, or catastrophic: Toward a refined taxonomy of overfitting
Neil Rohit Mallinar, James B Simon, Amirhesam Abedsoltan, Parthe Pandit, Misha Belkin, and Preetum Nakkiran · 2022
Later among the works it cites.
Harmless interpolation in regression and classification with structured features
Andrew D. Mcrae, Santhosh Karnik, Mark Davenport, and Vidya K. Muthukumar · 2022
Later among the works it cites.
Neural tangent kernel beyond the infinite-width limit: Effects of depth and initialization
Mariia Seleznova and Gitta Kutyniok · 2022
Later among the works it cites.
The implicit bias of benign overfitting
Ohad Shamir · 2022
Later among the works it cites.
Reverse engineering the neural tangent kernel
James Benjamin Simon, Sajant Anand, and Mike Deweese · 2022
Later among the works it cites.
Consistent interpolating ensembles via the manifold-Hilbert kernel
Yutong Wang and Clayton Scott · 2022
Later among the works it cites.
Strong inductive biases provably prevent harmless interpolation
Michael Aerni, Marco Milanta, Konstantin Donhauser, and Fanny Yang · 2023
Closest in time.
Is interpolation benign for random forest regression?
Ludovic Arnould, Claire Boyer, and Erwan Scornet · 2023
Closest in time.
On the inconsistency of kernel ridgeless regression in fixed dimensions
Daniel Beaglehole, Mikhail Belkin, and Parthe Pandit · 2023
Closest in time.
Sobolev spaces, kernels and discrepancies over hyperspheres
Simon Hubbert, Emilio Porcu, Chris J. Oates, and Mark Girolami · 2023
Closest in time.
mpmath: a Python library for arbitrary-precision floating-point arithmetic (version 1.3.0) , 2023
Fredrik Johansson et al · 2023
Closest in time.
Generalization ability of wide neural networks on ℝ \mathbb{R}
Jianfa Lai, Manyun Xu, Rui Chen, and Qian Lin · 2023
Closest in time.
Kernel interpolation generalizes poorly
Yicheng Li, Haobo Zhang, and Qian Lin · 2023
Closest in time.
Interpolating classifiers make few mistakes
Tengyuan Liang and Benjamin Recht · 2023
Closest in time.
For interpolating kernel machines, minimizing the norm of the erm solution maximizes stability
Akshay Rangamani, Lorenzo Rosasco, and Tomaso Poggio · 2023
Closest in time.
Benign overfitting of non-smooth neural networks beyond lazy training
Xingyu Xu and Yuantao Gu · 2023
Closest in time.