Fetching the paper…
Reading the bibliography…
Transformer language models (LMs) exhibit behaviors -- from storytelling to code generation -- that seem to require tracking the unobserved state of an evolving world.
Bounded-width polynomial-size branching programs recognize exactly those languages in NC1
Barrington, D. A · 1989
Earlier work this paper cites.
Does string-based neural MT learn source syntax?
Shi, X., Padhi, I., and Knight, K · 2016
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Bender, E. M. and Koller, A · 2020
Earlier work this paper cites.
On the Ability and Limitations of Transformers to Recognize Formal Languages
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Earlier work this paper cites.
BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance
McCoy, R. T., Min, J., and Linzen, T · 2020
Earlier work this paper cites.
Implicit Representations of Meaning in Neural Language Models
Li, B. Z., Nye, M., and Andreas, J · 2021
Earlier work this paper cites.
Locating and editing factual associations in GPT
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Earlier work this paper cites.
Saturated Transformers are Constant-Depth Threshold Circuits
Merrill, W., Sabharwal, A., and Smith, N. A · 2022
Earlier work this paper cites.
In-context Learning and Induction Heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Earlier work this paper cites.
Neural Networks and the Chomsky Hierarchy
Delétang, G., Ruoss, A., Grau-Moya, J., Genewein, T., Wenliang, L. K., Catt, E., Cundy, C., Hutter, M., Legg, S., Veness, J., and Ortega, P. A · 2023
Cited alongside, same era.
Latent state models of training dynamics
Hu, M. Y., Chen, A., Saphra, N., and Cho, K · 2023
Cited alongside, same era.
Entity Tracking in Language Models
Kim, N. and Schuster, S · 2023
Cited alongside, same era.
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M · 2023
Cited alongside, same era.
Transformers Learn Shortcuts to Automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2023
Cited alongside, same era.
The Parallelism Tradeoff: Limitations of Log-Precision Transformers
Merrill, W. and Sabharwal, A · 2023
Cited alongside, same era.
Emergent World Models and Latent Variable Estimation in Chess-Playing Language Models
Karvonen, A · 2024
Later among the works it cites.
A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformers
Merrill, W. and Sabharwal, A · 2024
Later among the works it cites.
The Illusion of State in State-Space Models
Merrill, W., Petty, J., and Sabharwal, A · 2024
Later among the works it cites.
Circuit Component Reuse Across Tasks in Transformer Language Models
Merullo, J., Eickhoff, C., and Pavlick, E · 2024
Later among the works it cites.
Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking
Prakash, N., Shaham, T. R., Haklay, T., Belinkov, Y., and Bau, D · 2024
Later among the works it cites.
What Formal Languages Can Transformers Express? A Survey
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emergent Linear Representations in World Models of Self-Supervised Sequence Models, 2023
Nanda, N., Lee, A., and Wattenberg, M · 2023
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2023
Cited alongside, same era.
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs
Chen, A., Shwartz-Ziv, R., Cho, K., Leavitt, M. L., and Saphra, N · 2024
Cited alongside, same era.
How to use and interpret activation patching
Heimersheim, S. and Nanda, N · 2024
Cited alongside, same era.
OthelloGPT learned a bag of heuristics, 2024
jylin04, JackS, Karvonen, A., and Can · 2024
Cited alongside, same era.
Strobl, L., Merrill, W., Weiss, G., Chiang, D., and Angluin, D · 2024
Later among the works it cites.
Evaluating the World Model Implicit in a Generative Model
Vafa, K., Chen, J. Y., Rambachan, A., Kleinberg, J., and Mullainathan, S · 2024
Later among the works it cites.
Do Large Language Models Latently Perform Multi-Hop Reasoning?
Yang, S., Gribovskaya, E., Kassner, N., Geva, M., and Riedel, S · 2024
Later among the works it cites.
Towards best practices of activation patching in language models: Metrics and methods
Zhang, F. and Nanda, N · 2024
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Zhong, Z., Liu, Z., Tegmark, M., and Andreas, J · 2024
Later among the works it cites.
Episodic memories generation and evaluation benchmark for large language models
Huet, A., Houidi, Z. B., and Rossi, D · 2025
Closest in time.