Fetching the paper…
Reading the bibliography…
Recent empirical studies show three phenomena with increasing size of language models: compute-optimal size scaling, emergent capabilities, and performance plateauing.
On Growth and Form
D’Arcy Wentworth Thompson · 1917
Earlier work this paper cites.
On being the right size
J. B. S. Haldane · 1926
Earlier work this paper cites.
A mathematical theory of communication
Claude E. Shannon · 1948
Earlier work this paper cites.
Exactly Solved Models in Statistical Mechanics
R. J. Baxter · 1982
Earlier work this paper cites.
SOAR: An architecture for general intelligence
John E. Laird, Allen Newell, and Paul S. Rosenbloom · 1987
Earlier work this paper cites.
Spin-glass models as error-correcting codes
Nicolas Sourlas · 1989
Earlier work this paper cites.
Unified Theories of Cognition
Allen Newell · 1990
Earlier work this paper cites.
Rules of the Mind
John R. Anderson · 1993
Earlier work this paper cites.
An overview of the EPIC architecture for cognition and performance with application to human-computer interaction
Davis E. Kieras and Davis E. Meyer · 1997
Earlier work this paper cites.
Finite-length analysis of low-density parity-check codes on the binary erasure channel
Changyan Di, David Proietti, I. Emre Telatar, Thomas J. Richardson, and Rüdiger L. Urbanke · 2002
Earlier work this paper cites.
The Theory of Information and Coding
Robert J. McEliece · 2002
Earlier work this paper cites.
Iterative quantization using codes on graphs
Emin Martinian and Jonathan S. Yedidia · 2003
Cited alongside, same era.
Modern Coding Theory
Tom Richardson and Rüdiger Urbanke · 2008
Cited alongside, same era.
Noise-enhanced associative memories
Amin Karbasi, Amir Hesam Salavati, Amin Shokrollahi, and Lav R. Varshney · 2013
Cited alongside, same era.
Network Science
Albert-László Barabási · 2016
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre · 2022
Are emergent abilities of large language models a mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2023
Later among the works it cites.
Skill-Mix: a flexible and expandable family of evaluations for AI models
Dingli Yu, Simran Kaur, Arushi Gupta, Jonah Brown-Cohen, Anirudh Goyal, and Sanjeev Arora · 2023
Later among the works it cites.
Information lattice learning
Haizi Yu, James A. Evans, and Lav R. Varshney · 2023
Later among the works it cites.
A group-theoretic approach to computational abstraction: Symmetry-driven hierarchical clustering
Haizi Yu, Igor Mineyev, and Lav R. Varshney · 2023
Later among the works it cites.
A dynamical model of neural scaling laws
Blake Bordelon, Alexander Atanasov, and Cengiz Pehlevan · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus · 2022
Cited alongside, same era.
A theory for emergence of complex skills in language models
Sanjeev Arora and Anirudh Goyal · 2023
Cited alongside, same era.
Transformers are universal predictors
Sourya Basu, Moulik Choraria, and Lav R. Varshney · 2023
Cited alongside, same era.
AI doom from an LLM-plateau-ist perspective, 2023
Steven Byrnes · 2023
Cited alongside, same era.
The quantization model of neural scaling
Eric Michaud, Ziming Liu, Uzay Girit, and Max Tegmark · 2023
Cited alongside, same era.
Anson Ho, Tamay Besiroglu, Ege Erdil, David Owen, Robi Rahman, Zifan Carl Guo, David Atkinson, Neil Thompson, and Jaime Sevilla · 2024
Closest in time.
On the limitations of compute thresholds as a governance strategy
Sara Hooker · 2024
Closest in time.
Unified view of grokking, double descent and emergent abilities: A comprehensive study on algorithm task
Yufei Huang, Shengding Hu, Xu Han, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
A mathematical theory for learning semantic languages by abstract learners
Kuo-Yu Liao, Cheng-Shang Chang, and Y.-W. Peter Hong · 2024
Closest in time.
The first wave of AI innovation is over. here’s what comes next
Gordon Ritter and Wendy Lu · 2024
Closest in time.