Fetching the paper…
Reading the bibliography…
In this work we investigate the generalization performance of random feature ridge regression (RFRR).
Gradient-based learning applied to document recognition
Yan Lecun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Kernels as features: On kernels, margins, and low-dimensional mappings
Maria-Florina Balcan, Avrim Blum, and Santosh Vempala · 2006
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Joel A Tropp et al · 2015
Earlier work this paper cites.
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
High-dimensional asymptotics of prediction: Ridge regression and classification
Edgar Dobriban and Stefan Wager · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
On the spectrum of random features maps of high dimensional data
Zhenyu Liao and Romain Couillet · 2018
Earlier work this paper cites.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J Zico Kolter, and Ryan J Tibshirani · 2019
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Limitations of lazy training of two-layers neural network
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
A jamming transition from under- to over-parametrization affects generalization in deep learning
S Spigler, M Geiger, S d’Ascoli, L Sagun, G Biroli, and M Wyart · 2019
Earlier work this paper cites.
The implicit regularization of stochastic gradient flow for least squares
Alnur Ali, Edgar Dobriban, and Ryan Tibshirani · 2020
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler · 2020
Earlier work this paper cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Earlier work this paper cites.
A precise performance analysis of learning with random features
Oussama Dhifallah and Yue M. Lu · 2020
Earlier work this paper cites.
Spectra of the conjugate kernel and neural tangent kernel for linear-width neural networks
Zhou Fan and Zhichao Wang · 2020
Earlier work this paper cites.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Cited alongside, same era.
On the optimal weighted \ell_2 regularization in overparameterized linear regression
Denny Wu and Ji Xu · 2020
Cited alongside, same era.
Deep learning: a statistical viewpoint
Peter L. Bartlett, Andrea Montanari, and Alexander Rakhlin · 2021
Cited alongside, same era.
Precise learning curves and higher-order scalings for dot-product kernel regression
Lechao Xiao, Hong Hu, Theodor Misiakiewicz, Yue Lu, and Jeffrey Pennington · 2022
Later among the works it cites.
What can be learnt with wide convolutional neural networks?
Francesco Cagnetta, Alessandro Favero, and Matthieu Wyart · 2023
Later among the works it cites.
Deterministic equivalent of the conjugate kernel matrix associated to artificial neural networks
Clément Chouard · 2023
Later among the works it cites.
Error scaling laws for kernel classification under source and capacity conditions
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2023
Later among the works it cites.
How two-layer neural networks learn, one (giant) step at a time
Yatin Dandi, Florent Krzakala, Bruno Loureiro, Luca Pesce, and Ludovic Stephan · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
Mikhail Belkin · 2021
Cited alongside, same era.
Locality defeats the curse of dimensionality in convolutional teacher-student scenarios
Alessandro Favero, Francesco Cagnetta, and Matthieu Wyart · 2021
Cited alongside, same era.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2021
Cited alongside, same era.
Generalization error of random feature and kernel methods: Hypercontractivity and kernel matrix concentration
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Cited alongside, same era.
Deep double descent: where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Cited alongside, same era.
Asymptotics of ridge(less) regression under general source condition
Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco · 2021
Cited alongside, same era.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Cited alongside, same era.
On the asymptotic learning curves of kernel ridge regression under power-law decay
Yicheng Li, haobo Zhang, and Qian Lin · 2023
Later among the works it cites.
Deterministic equivalent and error universality of deep random features learning
Dominik Schröder, Hugo Cui, Daniil Dmitriev, and Bruno Loureiro · 2023
Later among the works it cites.
Random features and polynomial rules
Fabián Aguirre-López, Silvio Franz, and Mauro Pastore · 2024
Closest in time.
Scaling and renormalization in high-dimensional regression
Alexander B. Atanasov, Jacob A. Zavatone-Veth, and Cengiz Pehlevan · 2024
Closest in time.
High-dimensional analysis of double descent for linear regression with random projections
Francis Bach · 2024
Closest in time.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2024
Closest in time.
A dynamical model of neural scaling laws
Blake Bordelon, Alexander Atanasov, and Cengiz Pehlevan · 2024
Closest in time.
Asymptotics of feature learning in two-layer networks after one gradient-step
Hugo Cui, Luca Pesce, Yatin Dandi, Florent Krzakala, Yue Lu, Lenka Zdeborova, and Bruno Loureiro · 2024
Closest in time.
Model collapse demystified: The case of regression
Elvis Dohmatob, Yunzhen Feng, and Julia Kempe · 2024
Closest in time.
Asymptotics of random feature regression beyond the linear scaling regime
Hong Hu, Yue M. Lu, and Theodor Misiakiewicz · 2024
Closest in time.
Scaling laws in linear regression: Compute, parameters, and data
Licong Lin, Jingfeng Wu, Sham M Kakade, Peter L Bartlett, and Jason D Lee · 2024
Closest in time.
Theodor Misiakiewicz and Basil Saeed · 2024
Closest in time.
A theory of non-linear feature learning with one gradient step in two-layer neural networks
Behrad Moniri, Donghwan Lee, Hamed Hassani, and Edgar Dobriban · 2024
Closest in time.
4+3 phases of compute-optimal neural scaling laws
Elliot Paquette, Courtney Paquette, Lechao Xiao, and Jeffrey Pennington · 2024
Closest in time.
Asymptotics of learning with deep structured (Random) features
Dominik Schröder, Daniil Dmitriev, Hugo Cui, and Bruno Loureiro · 2024
Closest in time.
On regularization via early stopping for least squares regression
Rishi Sonthalia, Jackie Lok, and Elizaveta Rebrova · 2024
Closest in time.
Nonlinear spiked covariance matrices and signal propagation in deep neural networks
Zhichao Wang, Denny Wu, and Zhou Fan · 2024
Closest in time.