Fetching the paper…
Reading the bibliography…
Empirically, large-scale deep learning models often satisfy a neural scaling law: the test error of the trained model improves polynomially as the model size and data size grow.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o ( 1 / n ) o(1/n)
Francis Bach and Eric Moulines · 2013
Earlier work this paper cites.
Sketching as a tool for numerical linear algebra
David P Woodruff et al · 2014
Earlier work this paper cites.
Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions
Alexandre Défossez and Francis Bach · 2015
Earlier work this paper cites.
Non-parametric stochastic approximation with large step sizes
Aymeric Dieuleveut and Francis R. Bach · 2015
Earlier work this paper cites.
Harder, better, faster, stronger convergence rates for least-squares regression
Aymeric Dieuleveut, Nicolas Flammarion, and Francis Bach · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification
Prateek Jain, Praneeth Netrapalli, Sham M Kakade, Rahul Kidambi, and Aaron Sidford · 2017
Earlier work this paper cites.
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco · 2017
Earlier work this paper cites.
Learning with sgd and random features
Luigi Carratino, Alessandro Rudi, and Lorenzo Rosasco · 2018
Earlier work this paper cites.
A Markov Chain Theory Approach to Characterizing the Minimax Optimality of Stochastic Gradient Descent (for Least Squares)
Prateek Jain, Sham M. Kakade, Rahul Kidambi, Praneeth Netrapalli, Venkata Krishna Pillutla, and Aaron Sidford · 2018
Earlier work this paper cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Earlier work this paper cites.
Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes
Loucas Pillaud-Vivien, Alessandro Rudi, and Francis Bach · 2018
Earlier work this paper cites.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares
Rong Ge, Sham M Kakade, Rahul Kidambi, and Praneeth Netrapalli · 2019
Earlier work this paper cites.
Aran Komatsuzaki · 2019
Earlier work this paper cites.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Martin J Wainwright · 2019
Cited alongside, same era.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Cited alongside, same era.
Tight nonparametric convergence rates for stochastic gradient descent under the noiseless linear model
Raphaël Berthier, Francis Bach, and Pierre Gaillard · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
A solvable model of neural scaling laws
Alexander Maloney, Daniel A Roberts, and James Sully · 2022
Later among the works it cites.
More than a toy: Random matrix models predict how real-world neural representations generalize
Alexander Wei, Wei Hu, and Jacob Steinhardt · 2022
Later among the works it cites.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Later among the works it cites.
Scaling data-constrained language models
Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel · 2023
Later among the works it cites.
Optimal eigenvalue approximation via sketching
William Swartworth and David P Woodruff · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
A neural scaling law from the dimension of the data manifold
Utkarsh Sharma and Jared Kaplan · 2020
Cited alongside, same era.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Cited alongside, same era.
Marcus Hutter · 2021
Cited alongside, same era.
Last iterate convergence of SGD for least-squares in the interpolation regime
Aditya Vardhan Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Cited alongside, same era.
The benefits of implicit regularization from sgd in least squares problems
Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, Dean P Foster, and Sham Kakade · 2021
Cited alongside, same era.
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Benign overfitting of constant-stepsize sgd for linear regression
Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, and Sham M Kakade · 2023
Later among the works it cites.
Scaling and renormalization in high-dimensional regression
Alexander B Atanasov, Jacob A Zavatone-Veth, and Cengiz Pehlevan · 2024
Closest in time.
Chinchilla scaling: A replication attempt
Tamay Besiroglu, Ege Erdil, Matthew Barnett, and Josh You · 2024
Closest in time.
A dynamical model of neural scaling laws
Blake Bordelon, Alexander Atanasov, and Cengiz Pehlevan · 2024
Closest in time.
Dimension-free deterministic equivalents for random feature regression
Leonardo Defilippis, Bruno Loureiro, and Theodor Misiakiewicz · 2024
Closest in time.
A tale of tails: Model collapse as a change of scaling laws
Elvis Dohmatob, Yunzhen Feng, Pu Yang, Francois Charton, and Julia Kempe · 2024
Closest in time.
Scaling laws for learning with real and surrogate data
Ayush Jain, Andrea Montanari, and Eren Sasoglu · 2024
Closest in time.
The quantization model of neural scaling
Eric Michaud, Ziming Liu, Uzay Girit, and Max Tegmark · 2024
Closest in time.
An exactly solvable model for emergence and scaling laws
Yoonsoo Nam, Nayara Fonseca, Seok Hyeong Lee, and Ard Louis · 2024
Closest in time.