Fetching the paper…
Reading the bibliography…
In recent years, significant attention in deep learning theory has been devoted to analyzing when models that interpolate their training data can still generalize well to unseen examples.
Spin Glass Theory and Beyond: An Introduction to the Replica Method and Its Applications
Marc Mézard, Giorgio Parisi, and Miguel Angel Virasoro · 1987
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Statistical Mechanics of Learning
Andreas Engel and Christian van den Broeck · 2001
Earlier work this paper cites.
Spectral moments of correlated Wishart matrices
Zdzisław Burda, Jerzy Jurkiewicz, and Bartłomiej Wacław · 2005
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
A. Caponnetto and E. De Vito · 2007
Earlier work this paper cites.
Aspects of Multivariate Statistical Theory
Robb J Muirhead · 2009
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2019
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler · 2020
Earlier work this paper cites.
Just interpolate: Kernel “Ridgeless” regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2020
Earlier work this paper cites.
Double trouble in double descent: Bias and variance(s) in the lazy regime
Stéphane d’Ascoli, Maria Refinetti, Giulio Biroli, and Florent Krzakala · 2020
Earlier work this paper cites.
Triple descent and the two kinds of overfitting: where & why do they appear?
Stéphane d’Ascoli, Levent Sagun, and Giulio Biroli · 2020
Earlier work this paper cites.
Understanding double descent requires a fine-grained bias-variance decomposition
Ben Adlam and Jeffrey Pennington · 2020
Earlier work this paper cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Earlier work this paper cites.
Finite versus infinite neural networks: an empirical study
Jaehoon Lee, Samuel Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Earlier work this paper cites.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mezard, and Lenka Zdeborova · 2020
Cited alongside, same era.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Cited alongside, same era.
Deep double descent: where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Cited alongside, same era.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Cited alongside, same era.
Statistical mechanics of deep linear neural networks: The backpropagating kernel renormalization
Qianyi Li and Haim Sompolinsky · 2021
Cited alongside, same era.
Minimum perturbation theory of deep perceptual learning
Haozhe Shan and Haim Sompolinsky · 2022
Later among the works it cites.
Structured random receptive fields enable informative sensory encodings
Biraj Pandey, Marius Pachitariu, Bingni W. Brunton, and Kameron Decker Harris · 2022
Later among the works it cites.
Asymptotics of representation learning in finite Bayesian neural networks
Jacob A Zavatone-Veth, Abdulkadir Canatar, Benjamin S Ruben, and Cengiz Pehlevan · 2022
Later among the works it cites.
Statistical mechanics of deep learning beyond the infinite-width limit
S. Ariosto, R. Pacelli, M. Pastore, F. Ginelli, M. Gherardi, and P. Rotondo · 2022
Later among the works it cites.
Exact learning dynamics of deep linear networks with prior knowledge
Lukas Braun, Clémentine Carla Juliette Dominé, James E Fitzgerald, and Andrew M Saxe · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning curves of generic features maps for realistic datasets with a teacher-student model
Bruno Loureiro, Cedric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mezard, and Lenka Zdeborova · 2021
Cited alongside, same era.
Out-of-distribution generalization in kernel regression
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2021
Cited alongside, same era.
Depth induces scale-averaging in overparameterized linear Bayesian neural networks
Jacob A Zavatone-Veth and Cengiz Pehlevan · 2021
Cited alongside, same era.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J. Tibshirani · 2022
Cited alongside, same era.
Universality laws for high-dimensional learning with random features
Hong Hu and Yue M. Lu · 2022
Cited alongside, same era.
Anisotropic random feature regression in high dimensions
Gabriel Mel and Jeffrey Pennington · 2022
Cited alongside, same era.
The Eigenlearning framework: A conservation law perspective on kernel regression and wide neural networks
James B. Simon, Madeline Dickens, Dhruva Karkada, and Michael R. DeWeese · 2022
Cited alongside, same era.
Deterministic equivalent and error universality of deep random features learning
Dominik Schröder, Hugo Cui, Daniil Dmitriev, and Bruno Loureiro · 2023
Closest in time.
Precise asymptotic analysis of deep random feature models
David Bosch, Ashkan Panahi, and Babak Hassibi · 2023
Closest in time.
The onset of variance-limited behavior for networks in the lazy and rich regimes
Alexander Atanasov, Blake Bordelon, Sabarish Sainathan, and Cengiz Pehlevan · 2023
Closest in time.
Bayes-optimal learning of deep random networks of extensive-width
Hugo Cui, Florent Krzakala, and Lenka Zdeborova · 2023
Closest in time.
Bayesian interpolation with deep linear networks
Boris Hanin and Alexander Zlokapa · 2023
Closest in time.
High-dimensional analysis of double descent for linear regression with random projections
Francis Bach · 2023
Closest in time.
Neural networks learn to magnify areas near decision boundaries
Jacob A. Zavatone-Veth, Sheng Yang, Julian A. Rubinfien, and Cengiz Pehlevan · 2023
Closest in time.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2041
Closest in time.
Strong replica symmetry for high-dimensional disordered log-concave Gibbs measures
Jean Barbier, Dmitry Panchenko, and Manuel Sáenz · 2049
Closest in time.