Fetching the paper…
Reading the bibliography…
Neural scaling laws characterize how model performance improves as the model size scales up.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Layer normalization, 2016
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
An introduction to systems biology: design principles of biological circuits
Uri Alon · 2019
Earlier work this paper cites.
An improved analysis of training over-parameterized deep neural networks, 2019
Difan Zou and Quanquan Gu · 2019
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
The depth-to-width interplay in self-attention
Yoav Levine, Noam Wies, Or Sharir, Hofit Bata, and Amnon Shashua · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
A neural scaling law from the dimension of the data manifold
Utkarsh Sharma and Jared Kaplan · 2020
Cited alongside, same era.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Cited alongside, same era.
Redundant representations help generalization in wide neural networks
Diego Doimo, Aldo Glielmo, Sebastian Goldt, and Alessandro Laio · 2021
Cited alongside, same era.
Data and parameter scaling laws for neural machine translation
Mitchell A Gordon, Kevin Duh, and Jared Kaplan · 2021
Cited alongside, same era.
Skill-it! a data-driven skills framework for understanding and training language models
Mayee F Chen, Nicholas Roberts, Kush Bhatia, Jue Wang, Ce Zhang, Frederic Sala, and Christopher Ré · 2023
Later among the works it cites.
The semantic landscape paradigm for neural networks
Shreyas Gokhale · 2023
Later among the works it cites.
Grokking modular arithmetic, 2023
Andrey Gromov · 2023
Later among the works it cites.
Finding neurons in a haystack: Case studies with sparse probing
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas · 2023
Later among the works it cites.
Transformers learn shortcuts to automata, 2023
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Do we actually need dense over-parameterization? in-time over-parameterization in sparse training
Shiwei Liu, Lu Yin, Decebal Constantin Mocanu, and Mykola Pechenizkiy · 2021
Cited alongside, same era.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine N Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso
Cited in the paper.
Towards automated circuit discovery for mechanistic interpretability, 2023b
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso
Cited in the paper.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah
Cited in the paper.
A neural scaling law from lottery ticket ensembling
Ziming Liu and Max Tegmark · 2023
Later among the works it cites.
The quantization model of neural scaling
Eric J Michaud, Ziming Liu, Uzay Girit, and Max Tegmark · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2023
Later among the works it cites.