Fetching the paper…
Reading the bibliography…
We derive simple closed-form estimates for the test risk and other generalization metrics of kernel ridge regression (KRR).
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 1901
Earlier work this paper cites.
Functions of positive and negative type and their connection with the theory of integral equations
J Mercer · 1909
Earlier work this paper cites.
Neural tangents: Fast and easy infinite neural networks in python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 1912
Earlier work this paper cites.
The lack of a priori distinctions between learning algorithms
David H Wolpert · 1996
Earlier work this paper cites.
Regression with gaussian processes: Average case performance
Manfred Opper · 1997
Earlier work this paper cites.
Learning curves for gaussian processes
Peter Sollich · 1999
Earlier work this paper cites.
Gaussian process regression with mismatched models
Peter Sollich · 2001
Earlier work this paper cites.
Learning curves for gaussian process regression: Approximations and bounds
Peter Sollich and Anason Halees · 2002
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2003
Earlier work this paper cites.
The curse of highly variable functions for local kernel machines
Yoshua Bengio, Olivier Delalleau, and Nicolas Le Roux · 2006
Earlier work this paper cites.
Roughness-induced critical phenomena in a turbulent flow
Nigel Goldenfeld · 2006
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
Statistical physics of fields
Mehran Kardar · 2007
Earlier work this paper cites.
The spectrum of kernel random matrices
Noureddine El Karoui · 2010
Earlier work this paper cites.
Spherical harmonics in p dimensions
Christopher Frye and Costas J Efthimiou · 2012
Earlier work this paper cites.
The spectrum of random inner-product kernel matrices
Xiuyuan Cheng and Amit Singer · 2013
Earlier work this paper cites.
Cavity method: Message passing from a physics perspective
Gino Del Ferraro, Chuang Wang, Dani Martí, and Marc Mézard · 2014
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
High-dimensional asymptotics of predictions: Ridge regression and classification
Edgar Dobriban and Stefan Wager · 2018
Earlier work this paper cites.
Justin Gilmer, Luke Metz, Fartash Faghri, Samuel S Schoenholz, Maithra Raghu, Martin Wattenberg, and Ian Goodfellow · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Clément Hongler, and Franck Gabriel · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle-Perez, Chico Q Camargo, and Ard A Louis · 2018
Earlier work this paper cites.
Understanding training and generalization in deep learning by fourier analysis
Zhiqin John Xu · 2018
Cited alongside, same era.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J Zico Kolter, and Ryan J Tibshirani · 2019
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
Towards understanding the spectral bias of deep learning
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu · 2019
Cited alongside, same era.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Later among the works it cites.
On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression
Denny Wu and Ji Xu · 2020
Later among the works it cites.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Closest in time.
Deep learning: a statistical viewpoint
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin · 2021
Closest in time.
Learning curves for sgd on structured features
Blake Bordelon and Cengiz Pehlevan · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The spectral norm of random inner-product kernel matrices
Zhou Fan and Andrea Montanari · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
On learning over-parameterized neural networks: A functional approximation perspective
Lili Su and Pengkun Yang · 2019
Cited alongside, same era.
A fine-grained spectral perspective on neural networks
Greg Yang and Hadi Salman · 2019
Cited alongside, same era.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2021
Closest in time.
Learning curves for overparametrized deep neural networks: A field theory perspective
Omry Cohen, Or Malka, and Zohar Ringel · 2021
Closest in time.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2021
Closest in time.
Elvis Dohmatob · 2021
Closest in time.
Unsolved problems in ml safety
Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt · 2021
Closest in time.
Dimension lower bounds for linear approaches to function approximation
Daniel Hsu · 2021
Closest in time.
Kernel regression in high dimensions: Refined analysis beyond double descent
Fanghui Liu, Zhenyu Liao, and Johan Suykens · 2021
Closest in time.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Bruno Loureiro, Cedric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mezard, and Lenka Zdeborová · 2021
Closest in time.
Learning with convolution and pooling operations in kernel methods
Theodor Misiakiewicz and Song Mei · 2021
Closest in time.
Optimal regularization can mitigate double descent
Preetum Nakkiran, Prayaag Venkat, Sham M. Kakade, and Tengyu Ma · 2021
Closest in time.
Asymptotics of ridge (less) regression under general source condition
Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco · 2021
Closest in time.
Eigenspace restructuring: a principle of space and frequency in neural networks
Lechao Xiao · 2021
Closest in time.
How wide convolutional neural networks learn hierarchical tasks
Francesco Cagnetta, Alessandro Favero, and Matthieu Wyart · 2022
Closest in time.
The three stages of learning dynamics in high-dimensional kernel methods
Nikhil Ghosh, Song Mei, and Bin Yu · 2022
Closest in time.
An equivalence principle for the spectrum of random inner-product kernel matrices
Yue M Lu and Horng-Tzer Yau · 2022
Closest in time.
Benign, tempered, or catastrophic: A taxonomy of overfitting
Neil Mallinar, James B Simon, Amirhesam Abedsoltan, Parthe Pandit, Mikhail Belkin, and Preetum Nakkiran · 2022
Closest in time.
Failure and success of the spectral bias prediction for laplace kernel ridge regression: the case of low-dimensional data
Umberto M Tomasini, Antonio Sclocchi, and Matthieu Wyart · 2022
Closest in time.
More than a toy: Random matrix models predict how real-world neural representations generalize
Alexander Wei, Wei Hu, and Jacob Steinhardt · 2022
Closest in time.