Fetching the paper…
Reading the bibliography…
Neural scaling laws have garnered significant interest due to their ability to predict model performance as a function of increasing parameters, data, and compute.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Scaling laws for neural language models, 2020b
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
Optimal rates for regularized least squares regression
Ingo Steinwart, Don R Hush, Clint Scovel, et al · 2009
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2010
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling, 2020b
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B. Brown, Prafulla Dhariwal, Scott Gray, Chris Hallacy, Benjamin Mann, Alec Radford, Aditya Ramesh, Nick Ryder, Daniel M. Ziegler, John Schulman, Dario Amodei, and Sam McCandlish · 2010
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Critical percolation as a framework to analyze the training of deep networks, 2018
Zohar Ringel and Rodrigo de Bem · 2018
Earlier work this paper cites.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Cited alongside, same era.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Cited alongside, same era.
Sobolev norm learning rates for regularized least-squares algorithms
Simon Fischer and Ingo Steinwart · 2020
Cited alongside, same era.
Marcus Hutter · 2021
Cited alongside, same era.
The universal statistical structure and scaling laws of chaos and turbulence, 2023
Noam Levi and Yaron Oz · 2023
Later among the works it cites.
Efficiently programming large language models using sglang, 2023
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng · 2023
Later among the works it cites.
Large language monkeys: Scaling inference compute with repeated sampling, 2024
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christopher Ré, and Azalia Mirhoseini · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Closest in time.
Smaller, weaker, yet better: Training llm reasoners via compute-optimal sampling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2021
Cited alongside, same era.
Mostafa Dehghani, Anurag Arnab, Lucas Beyer, Ashish Vaswani, and Yi Tay · 2021
Cited alongside, same era.
An empirical analysis of compute-optimal large language model training
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
A solvable model of neural scaling laws
Alexander Maloney, Daniel A Roberts, and James Sully · 2022
Cited alongside, same era.
Hritik Bansal, Arian Hosseini, Rishabh Agarwal, Vinh Q. Tran, and Mehran Kazemi · 2024
Closest in time.
Scaling laws for associative memories, 2024
Vivien Cabannes, Elvis Dohmatob, and Alberto Bietti · 2024
Closest in time.
The underlying scaling laws and universal statistical structure of complex datasets, 2024
Noam Levi and Yaron Oz · 2024
Closest in time.
Yoonsoo Nam, Nayara Fonseca, Seok Hyeong Lee, Chris Mingard, and Ard A. Louis · 2024
Closest in time.
Hydragen: High-throughput llm inference with shared prefixes
Jordan Juravsky, Bradley Brown, Ryan Ehrlich, Daniel Y. Fu, Christopher R’e, and Azalia Mirhoseini · 2024
Closest in time.