Fetching the paper…
Reading the bibliography…
For a given distribution, learning algorithm, and performance metric, the rate of convergence (or data-scaling law) is the asymptotic behavior of the algorithm's test performance as a function of number of train samples.
A formal theory of inductive inference. part i
Ray J. Solomonoff · 1964
Earlier work this paper cites.
Nearest neighbor pattern classification
Thomas Cover and Peter Hart · 1967
Earlier work this paper cites.
Universal problems of full search
Leonid A. Levin · 1973
Earlier work this paper cites.
Complexity-based induction systems: comparisons and convergence theorems
Ray J. Solomonoff · 1978
Earlier work this paper cites.
Any discrimination rule can have an arbitrarily bad probability of error for finite sample size
Luc Devroye · 1982
Earlier work this paper cites.
One-way functions and pseudorandom generators
Leonid A. Levin · 1985
Earlier work this paper cites.
Training connectionist networks with queries and selective sampling
Les E Atlas, David A Cohn, and Richard E Ladner · 1990
Earlier work this paper cites.
Learning from examples in large neural networks
Haim Sompolinsky, Naftali Tishby, and H Sebastian Seung · 1990
Earlier work this paper cites.
Statistical Theory of Learning Curves under Entropic Loss Criterion
Shun-ichi Amari and Noboru Murata · 1993
Earlier work this paper cites.
The lack of a priori distinctions between learning algorithms
David H Wolpert · 1996
Earlier work this paper cites.
No free lunch theorems for optimization
David H Wolpert and William G Macready · 1997
Cited alongside, same era.
A theory of universal artificial intelligence based on algorithmic complexity
Marcus Hutter · 2000
Cited alongside, same era.
Tree induction vs. logistic regression: A learning-curve analysis
Claudia Perlich, Foster Provost, and Jeffrey Simonoff · 2003
Cited alongside, same era.
Universal Artificial Intelligence: Sequential Decisions based on Algorithmic Probability
Marcus Hutter · 2005
Cited alongside, same era.
Introduction to nonparametric estimation., 2009
Alexandre B Tsybakov · 2009
Cited alongside, same era.
Introduction to online optimization
Sébastien Bubeck · 2011
Cited alongside, same era.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou · 2017
Cited alongside, same era.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Cited alongside, same era.
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Later among the works it cites.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Closest in time.
A theory of universal learning
Olivier Bousquet, Steve Hanneke, Shay Moran, Ramon van Handel, and Amir Yehudayoff · 2021
Closest in time.
Marcus Hutter · 2021
Closest in time.
The shape of learning curves: a review
Tom Viering and Marco Loog · 2021
Closest in time.