Fetching the paper…
Reading the bibliography…
What computational structure are we building into large language models when we train them on next-token prediction? Here, we present evidence that this structure is given by the meta-dynamics of belief updating over hidden states of the data-generating process.
Theory and Algorithms for Hidden Markov Models and Generalized Hidden Markov Models
D. R. Upper · 1997
Earlier work this paper cites.
Computational mechanics: Pattern and prediction, structure and simplicity
C. R. Shalizi and J. P. Crutchfield · 2001
Earlier work this paper cites.
Synchronization and control in intrinsic and designed computation: An information-theoretic analysis of competing models of stochastic computation
J. P. Crutchfield, C. J. Ellison, J. R. Mahoney, and R. G. James · 2010
Earlier work this paper cites.
Between order and chaos
J. P. Crutchfield · 2012
Earlier work this paper cites.
Predictive models and generative complexity
Wolfgang Löhr · 2012
Earlier work this paper cites.
Signatures of infinity: Nonergodicity and resource scaling in prediction, complexity, and learning
James P Crutchfield and Sarah Marzen · 2015
Earlier work this paper cites.
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Prediction and generation of binary markov processes: Can a finite-state fox catch a markov mouse?
Joshua B Ruebeck, Ryan G James, John R Mahoney, and James P Crutchfield · 2018
Earlier work this paper cites.
Meta-learning of sequential strategies
Pedro A Ortega, Jane X Wang, Mark Rowland, Tim Genewein, Zeb Kurth-Nelson, Razvan Pascanu, Nicolas Heess, Joel Veness, Alex Pritzel, Pablo Sprechmann, et al · 2019
Earlier work this paper cites.
Information theory meets power laws: Stochastic processes and language models
Lukasz Debowski · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Cited alongside, same era.
Fraudulent white noise: Flat power spectra belie arbitrarily complex processes
Paul M Riechers and James P Crutchfield · 2021
Cited alongside, same era.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2022
Cited alongside, same era.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2022
Cited alongside, same era.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
Future lens: Anticipating subsequent tokens from a single hidden state
Koyena Pal, Jiuding Sun, Andrew Yuan, Byron C Wallace, and David Bau · 2023
Later among the works it cites.
The pitfalls of next-token prediction
Gregor Bachmann and Vaishnavh Nagarajan · 2024
Closest in time.
Not all language model features are linear
Joshua Engels, Isaac Liao, Eric J Michaud, Wes Gurnee, and Max Tegmark · 2024
Closest in time.
Better & faster large language models via multi-token prediction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Transformerlens
Neel Nanda and Joseph Bloom · 2022
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Cited alongside, same era.
A simplistic model of neural scaling laws: Multiperiodic santa fe processes
Łukasz Dębowski · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Cited alongside, same era.
Shannon entropy rate of hidden Markov processes
A. M. Jurgens and J. P. Crutchfield
Cited in the paper.
Divergent predictive states: The statistical complexity dimension of stationary, ergodic hidden markov processes
Alexandra M Jurgens and James P Crutchfield
Cited in the paper.
Nearly maximally predictive features and their dimensions
S. E. Marzen and J. P. Crutchfield
Cited in the paper.
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, and Gabriel Synnaeve · 2024
Closest in time.
Handcrafting a network to predict next token probabilities for the random-random-xor process
Rick Goldstein · 2024
Closest in time.
On limitation of transformer for learning hmms
Jiachen Hu, Qinghua Liu, and Chi Jin · 2024
Closest in time.
RNNs represent belief state geometry in their hidden states
Keenan Pepper · 2024
Closest in time.
Building AGI, alignment, future models, spies, Microsoft, Taiwan, & enlightenment
Ilya Sutskever · 2024
Closest in time.