Fetching the paper…
Reading the bibliography…
It is by now well-established that modern over-parameterized models seem to elude the bias-variance tradeoff and generalize well despite overfitting noise.
Orthogonal polynomials , volume 23
Gabor Szeg · 1939
Earlier work this paper cites.
A quantitative formulation of sylvester’s law of inertia. iii
Jerome Dancis · 1986
Earlier work this paper cites.
Regularization with dot-product kernels
Alex Smola, Zoltán Ovári, and Robert C Williamson · 2000
Earlier work this paper cites.
Spectral properties of the kernel matrix and their relation to kernel methods in machine learning
Mikio Ludwig Braun · 2005
Earlier work this paper cites.
Fast rates for regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2005
Earlier work this paper cites.
Mercer’s theorem, feature maps, and smoothing
Ha Quang Minh, Partha Niyogi, and Yuan Yao · 2006
Earlier work this paper cites.
An explicit description of the reproducing kernel hilbert spaces of gaussian rbf kernels
Ingo Steinwart, Don Hush, and Clint Scovel · 2006
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
Optimal rates for regularized least squares regression
Ingo Steinwart, Don R Hush, Clint Scovel, et al · 2009
Earlier work this paper cites.
On learning with integral operators
Lorenzo Rosasco, Mikhail Belkin, and Ernesto De Vito · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Positive definite kernels: past, present and future
Gregory E Fasshauer · 2011
Earlier work this paper cites.
Spherical harmonics and approximations on the unit sphere: an introduction , volume 2044
Kendall Atkinson and Weimin Han · 2012
Earlier work this paper cites.
Matrix analysis
Roger A Horn and Charles R Johnson · 2012
Earlier work this paper cites.
Mercer’s theorem on general domains: On the interaction between measures, kernels, and rkhss
Ingo Steinwart and Clint Scovel · 2012
Earlier work this paper cites.
Approximation theory and harmonic analysis on spheres and balls
Feng Dai · 2013
Earlier work this paper cites.
Analysis of boolean functions
Ryan O’Donnell · 2014
Earlier work this paper cites.
Eigenvalues of dot-product kernels on the sphere
Douglas Azevedo and Valdir A Menegatto · 2015
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Joel A Tropp et al · 2015
Earlier work this paper cites.
Spherical harmonics with maximal l p (2 < < p ≤ \leq 6) norm growth
Xiaolong Han · 2016
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Approximation beats concentration? an approximation view on inference with smooth radial kernels
Mikhail Belkin · 2018
Earlier work this paper cites.
High-dimensional asymptotics of prediction: Ridge regression and classification
Edgar Dobriban and Stefan Wager · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Relative concentration bounds for the spectrum of kernel matrices
Ernesto Araya Valdivia · 2018
Earlier work this paper cites.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J Zico Kolter, and Ryan J Tibshirani · 2019
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
The convergence rate of neural networks for learned functions of different frequencies
Ronen Basri, David Jacobs, Yoni Kasten, and Shira Kritchman · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
Towards understanding the spectral bias of deep learning
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu · 2019
Cited alongside, same era.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Later among the works it cites.
Towards an understanding of benign overfitting in neural networks
Zhu Li, Zhi-Hua Zhou, and Arthur Gretton · 2021
Later among the works it cites.
Tight bounds on the smallest eigenvalue of the neural tangent kernel for deep relu networks
Quynh Nguyen, Marco Mondelli, and Guido F Montufar · 2021
Later among the works it cites.
Asymptotics of ridge (less) regression under general source condition
Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco · 2021
Later among the works it cites.
A spectral analysis of dot-product kernels
Meyer Scetbon and Zaid Harchaoui · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Consistency of interpolation with laplace kernels is a high-dimensional phenomenon
Alexander Rakhlin and Xiyu Zhai · 2019
Cited alongside, same era.
A fine-grained spectral perspective on neural networks
Greg Yang and Hadi Salman · 2019
Cited alongside, same era.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Cited alongside, same era.
Deep equals shallow for relu networks in kernel regimes
Alberto Bietti and Francis Bach · 2020
Cited alongside, same era.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Cited alongside, same era.
James B Simon, Madeline Dickens, Dhruva Karkada, and Michael R DeWeese · 2021
Later among the works it cites.
Zhichao Wang and Yizhe Zhu · 2021
Later among the works it cites.
Tensor programs iib: Architectural universality of neural tangent kernel training dynamics
Greg Yang and Etai Littwin · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
A kernel perspective of skip connections in convolutional networks
Daniel Barzilai, Amnon Geifman, Meirav Galun, and Ronen Basri · 2022
Later among the works it cites.
Spectral bias outside the training set for deep networks in the kernel regime
Benjamin Bowman and Guido F Montufar · 2022
Later among the works it cites.
Dimension free ridge regression
Chen Cheng and Andrea Montanari · 2022
Later among the works it cites.
On the spectral bias of convolutional neural tangent and gaussian process kernels
Amnon Geifman, Meirav Galun, David Jacobs, and Basri Ronen · 2022
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2022
Later among the works it cites.
Universality laws for high-dimensional learning with random features
Hong Hu and Yue M Lu · 2022
Later among the works it cites.
Benign, tempered, or catastrophic: A taxonomy of overfitting
Neil Mallinar, James B Simon, Amirhesam Abedsoltan, Parthe Pandit, Mikhail Belkin, and Preetum Nakkiran · 2022
Later among the works it cites.
Generalization error of random feature and kernel methods: hypercontractivity and kernel matrix concentration
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2022
Later among the works it cites.
Theodor Misiakiewicz · 2022
Later among the works it cites.
The interpolation phase transition in neural networks: Memorization and generalization under lazy training
Andrea Montanari and Yiqiao Zhong · 2022
Later among the works it cites.
Precise learning curves and higher-order scaling limits for dot product kernel regression
Lechao Xiao and Jeffrey Pennington · 2022
Later among the works it cites.
High-dimensional analysis of double descent for linear regression with random projections
Francis Bach · 2023
Closest in time.
A theoretical analysis of the test error of finite-rank kernel ridge regression
Tin Sum Cheng, Aurelien Lucchi, Ivan Dokmanić, Anastasis Kratsios, and David Belius · 2023
Closest in time.
Controlling the inductive bias of wide neural networks by modifying the kernel’s spectrum
Amnon Geifman, Daniel Barzilai, Ronen Basri, and Meirav Galun · 2023
Closest in time.
Mind the spikes: Benign overfitting of kernels and neural networks in fixed dimension
Moritz Haas, David Holzmüller, Ulrike von Luxburg, and Ingo Steinwart · 2023
Closest in time.
Benign overfitting in ridge regression
Alexander Tsigler and Peter L Bartlett · 2023
Closest in time.