Fetching the paper…
Reading the bibliography…
Large language models with a huge number of parameters, when trained on near internet-sized number of tokens, have been empirically shown to obey neural scaling laws: specifically, their performance behaves predictably as a power law in either parameters or dataset size until bottlenecked by the other resource.
K. Pearson, “On lines and planes of closest fit to systems of points in space,” Philos. Mag
1901
Earlier work this paper cites.
1905
Earlier work this paper cites.
1910
Earlier work this paper cites.
1912
Earlier work this paper cites.
F. J. Dyson, “The S S Matrix in Quantum Electrodynamics,” Phys. Rev
1949
Earlier work this paper cites.
J. Schwinger, “On the Green’s functions of quantized fields. I,” Proceedings of the National Academy of Sciences
1951
Earlier work this paper cites.
V. A. Marčenko and L. A. Pastur, “Distribution of Eigenvalues for Some Sets of Random Matrices,” Mathematics of the USSR-Sbornik
1967
Earlier work this paper cites.
G. ’t Hooft, “A Planar Diagram Theory for Strong Interactions,” Nucl. Phys
1974
Earlier work this paper cites.
J. W. Silverstein and Z. D. Bai, “On the Empirical Distribution of Eigenvalues of a Class of Large Dimensional Random Matrices,” Journal of Multivariate Analysis
1995
Earlier work this paper cites.
M. Opper, “Statistical mechanics of learning: Generalization,” The handbook of brain theory and neural networks
1995
Earlier work this paper cites.
Springer, 1996
R. M. Neal, “Priors for infinite networks,” in Bayesian Learning for Neural Networks · 1996
Earlier work this paper cites.
2001
Earlier work this paper cites.
M. Opper, “Learning to generalize,” Frontiers of Life
2001
Earlier work this paper cites.
https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.3.4934&rep=rep1&type=pdf
J. Friedman and B. E. Popescu, “Gradient directed regularization for linear regression and classification,” tech. rep., 2002 · 2002
Earlier work this paper cites.
2004
Cited alongside, same era.
M. Mitzenmacher, “A brief history of generative models for power law and lognormal distributions,” Internet mathematics
2004
Cited alongside, same era.
Z. Burda, A. Görlich, A. Jarosz, and J. Jurkiewicz, “Signal and noise in correlation matrix,” Physica A: Statistical Mechanics and its Applications
2004
Cited alongside, same era.
MIT Press, Cambridge, MA, USA, 2004
E. Levina and P. J. Bickel, “Maximum Likelihood Estimation of Intrinsic Dimension,” in Proceedings of the 17th International Conference on Neural Information Processing Systems · 2004
Cited alongside, same era.
S. Bradde and W. Bialek, “PCA meets RG,” Journal of Statistical Physics
2017
Later among the works it cites.
Springer, 2017
J. A. Mingo and R. Speicher, Free Probability and Random Matrices , vol. 35 · 2017
Later among the works it cites.
E. Facco, M. d’Errico, A. Rodriguez, and A. Laio, “Estimating the intrinsic dimension of datasets by a minimal neighborhood information,” Scientific Reports
2017
Later among the works it cites.
2018
Later among the works it cites.
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2005
Cited alongside, same era.
2006
Cited alongside, same era.
Springer Science & Business Media, 2009
T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning · 2009
Cited alongside, same era.
2009
Cited alongside, same era.
https://www.cs.toronto.edu/~kriz/cifar.html
A. Krizhevsky, “Learning multiple layers of features from tiny images,” 2009 · 2009
Cited alongside, same era.
2010
Cited alongside, same era.
2010
Cited alongside, same era.
2014
Cited alongside, same era.
Later among the works it cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” OpenAI Blog
2019
Later among the works it cites.
B. Adlam, J. Levinson, and J. Pennington, “A Random Matrix Perspective on Mixtures of Nonlinearities for Deep Learning,” 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
https://www.gwern.net/Scaling-hypothesis
G. Branwen, “The Scaling Hypothesis,” May, 2020 · 2020
Later among the works it cites.
2020
Later among the works it cites.
Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, Nov., 2021
M. A. Gordon, K. Duh, and J. Kaplan, “Data and Parameter Scaling Laws for Neural Machine Translation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing · 2021
Later among the works it cites.
https://www.alignmentforum.org/posts/GzoWcYibWYwJva8aL/parameter-counts-in-machine-learning
J. Sevilla and P. Villalobos, “Parameter counts in Machine Learning,” Jun, 2021 · 2021
Later among the works it cites.
Cambridge University Press, 2022
D. A. Roberts, S. Yaida, and B. Hanin, The Principles of Deep Learning Theory · 2022
Closest in time.