Fetching the paper…
Reading the bibliography…
Neural scaling laws (NSL) refer to the phenomenon where model performance improves with scale.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Grant Rotskoff and Eric Vanden-Eijnden · 2018
Earlier work this paper cites.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Earlier work this paper cites.
A neural scaling law from the dimension of the data manifold
Utkarsh Sharma and Jared Kaplan · 2020
Earlier work this paper cites.
The depth-to-width interplay in self-attention
Yoav Levine, Noam Wies, Or Sharir, Hofit Bata, and Amnon Shashua · 2020
Cited alongside, same era.
On power laws in deep ensembles
Ekaterina Lobacheva, Nadezhda Chirkova, Maxim Kodryan, and Dmitry P Vetrov · 2020
Cited alongside, same era.
Data and parameter scaling laws for neural machine translation
Mitchell A Gordon, Kevin Duh, and Jared Kaplan · 2021
Cited alongside, same era.
Redundant representations help generalization in wide neural networks
Diego Doimo, Aldo Glielmo, Sebastian Goldt, and Alessandro Laio · 2021
Cited alongside, same era.
Neural networks and quantum field theory
James Halverson, Anindita Maiti, and Keegan Stoner · 2021
Cited alongside, same era.
Scaling vision transformers
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Later among the works it cites.
The principles of deep learning theory
Daniel A Roberts, Sho Yaida, and Boris Hanin · 2022
Later among the works it cites.
Are emergent abilities of large language models a mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2023
Closest in time.
The hydra effect: Emergent self-repair in language model computations
Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg · 2023
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
The quantization model of neural scaling
Eric J Michaud, Ziming Liu, Uzay Girit, and Max Tegmark
Cited in the paper.
Precision machine learning
Eric J Michaud, Ziming Liu, and Max Tegmark
Cited in the paper.
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2023
Closest in time.
Seeing is believing: Brain-inspired modular training for mechanistic interpretability
Ziming Liu, Eric Gan, and Max Tegmark · 2023
Closest in time.