Fetching the paper…
Reading the bibliography…
We consider the solvable neural scaling model with three parameters: data complexity, target complexity, and model-parameter-count.
An Elementary Proof of Error Estimates for the Trapezoidal Rule
D. Cruz-Uribe and C. J. Neugebauer · 1930
Earlier work this paper cites.
On the resolvents of nonconvolution Volterra kernels
Gustaf Gripenberg · 1980
Earlier work this paper cites.
On the empirical distribution of eigenvalues of a class of large dimensional random matrices
Jack W Silverstein and Zhi Dong Bai · 1995
Earlier work this paper cites.
Applied probability and queues , volume 2
Søren Asmussen, Soren Asmussen, and Sren Asmussen · 2003
Earlier work this paper cites.
Branching processes
Krishna B Athreya, Peter E Ney, and PE Ney · 2004
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
Deterministic equivalents for certain functionals of large random matrices
Walid Hachem, Philippe Loubaton, and Jamal Najim · 2007
Earlier work this paper cites.
Spectral analysis of large dimensional random matrices , volume 20
Zhidong Bai and Jack W Silverstein · 2010
Earlier work this paper cites.
Optimal Distributed Online Prediction Using Mini-Batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Earlier work this paper cites.
Nonparametric stochastic approximation with large step-sizes
Aymeric Dieuleveut and Francis Bach · 2016
Earlier work this paper cites.
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco · 2017
Earlier work this paper cites.
Learning with sgd and random features
Luigi Carratino, Alessandro Rudi, and Lorenzo Rosasco · 2018
Earlier work this paper cites.
Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes
Loucas Pillaud-Vivien, Alessandro Rudi, and Francis Bach · 2018
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alex Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
A neural scaling law from the dimension of the data manifold
Utkarsha Sharma and Jared Kaplan · 2020
Cited alongside, same era.
Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy Regime
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2021
Cited alongside, same era.
On the interplay between data structure and loss function in classification problems
Stéphane d’Ascoli, Marylou Gabrié, Levent Sagun, and Giulio Biroli · 2021
Cited alongside, same era.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Bruno Loureiro, Cedric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mezard, and Lenka Zdeborová · 2021
Cited alongside, same era.
Anisotropic random feature regression in high dimensions
Gabriel Mel and Jeffrey Pennington · 2021
Cited alongside, same era.
Homogenization of SGD in high-dimensions: exact dynamics and generalization properties
Courtney Paquette, Elliot Paquette, Ben Adlam, and Jeffrey Pennington · 2022
Later among the works it cites.
From high-dimensional & mean-field dynamics to dimensionless odes: A unifying approach to sgd in two-layers networks
Luca Arnaboldi, Ludovic Stephan, Florent Krzakala, and Bruno Loureiro · 2023
Later among the works it cites.
Benign overfitting of constant-stepsize SGD for linear regression
Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, and Sham M Kakade · 2023
Later among the works it cites.
High-dimensional analysis of double descent for linear regression with random projections
Francis Bach · 2024
Closest in time.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamics of stochastic momentum methods on large-scale, quadratic models
Courtney Paquette and Elliot Paquette · 2021
Cited alongside, same era.
SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality
Courtney Paquette, Kiwon Lee, Fabian Pedregosa, and Elliot Paquette · 2021
Cited alongside, same era.
Last iterate convergence of SGD for Least-Squares in the Interpolation regime
Aditya Vardhan Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Cited alongside, same era.
Learning Curves for SGD on Structured Features
Blake Bordelon and Cengiz Pehlevan · 2022
Cited alongside, same era.
Dimension free ridge regression
Chen Cheng and Andrea Montanari · 2022
Cited alongside, same era.
Random Matrix Methods for Machine Learning
Romain Couillet and Zhenyu Liao · 2022
Cited alongside, same era.
An empirical analysis of compute-optimal large language model training
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack Rae, and Laurent Sifre · 2022
Cited alongside, same era.
Tamay Besiroglu, Ege Erdil, Matthew Barnett, and Josh You · 2024
Closest in time.
A Dynamical Model of Neural Scaling Laws
Blake Bordelon, Alexander Atanasov, and Cengiz Pehlevan · 2024
Closest in time.
Dimension-free deterministic equivalents for random feature regression
Leonardo Defilippis, Bruno Loureiro, and Theodor Misiakiewicz · 2024
Closest in time.
Rigorous dynamical mean-field theory for stochastic gradient descent methods
Cedric Gerbelot, Emanuele Troiani, Francesca Mignacco, Florent Krzakala, and Lenka Zdeborova · 2024
Closest in time.
Asymptotics of Random Feature Regression Beyond the Linear Scaling Regime
Hong Hu, Yue M. Lu, and Theodor Misiakiewicz · 2024
Closest in time.
Scaling Laws in Linear Regression: Compute, Parameters, and Data
Licong Lin, Jingfeng Wu, Sham M Kakade, Peter L Bartlett, and Jason D Lee · 2024
Closest in time.
A Solvable Model of Neural Scaling Laws
Alexander Maloney, Daniel A. Roberts, and James Sully · 2024
Closest in time.
More is better in modern machine learning: when infinite overparameterization is optimal and overfitting is obligatory
James B. Simon, Dhruva Karkada, Nikhil Ghosh, and Mikhail Belkin · 2024
Closest in time.
Hitting the high-dimensional notes: an ODE for SGD learning dynamics on GLMs and multi-index models
Elizabeth Collins-Woodfin, Courtney Paquette, Elliot Paquette, and Inbar Seroussi · 2049
Closest in time.