Fetching the paper…
Reading the bibliography…
We investigate the internal structure of language model computations using causal analysis and demonstrate two motifs: (1) a form of adaptive computation where ablations of one attention layer of a language model cause another layer to compensate (which we term the Hydra effect) and (2) a counterbalancing function of late MLP layers that act to downregulate the maximum-likelihood token.
On aims and methods of ethology
N. Tinbergen · 1963
Earlier work this paper cites.
Causal mediation analysis for interpreting neural nlp: The case of gender bias
J. Vig, S. Gehrmann, Y. Belinkov, S. Qian, D. Nevo, S. Sakenis, J. Huang, Y. Singer, and S. Shieber · 2004
Earlier work this paper cites.
Causality
J. Pearl · 2009
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
G. Alain and Y. Bengio · 2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Causal inference in statistics: A primer
M. Glymour, J. Pearl, and N. P. Jewell · 2016
Earlier work this paper cites.
Actual causality
J. Y. Halpern · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
A. Veit, M. J. Wilber, and S. Belongie · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Gan dissection: Visualizing and understanding generative adversarial networks
D. Bau, J.-Y. Zhu, H. Strobelt, B. Zhou, J. B. Tenenbaum, W. T. Freeman, and A. Torralba · 2018
Earlier work this paper cites.
On the importance of single directions for generalization
A. S. Morcos, D. G. Barrett, N. C. Rabinowitz, and M. Botvinick · 2018
Earlier work this paper cites.
The building blocks of interpretability
C. Olah, A. Satyanarayan, I. Johnson, S. Carter, L. Schubert, K. Ye, and A. Mordvintsev · 2018
Earlier work this paper cites.
Activation atlas
S. Carter, Z. Armstrong, L. Schubert, I. Johnson, and C. Olah · 2019
Earlier work this paper cites.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Cited alongside, same era.
Curve detectors
N. Cammarata, G. Goh, S. Carter, L. Schubert, M. Petrov, and C. Olah · 2020
Cited alongside, same era.
Understanding rl vision
J. Hilton, N. Cammarata, S. Carter, G. Goh, and C. Olah · 2020
Cited alongside, same era.
Towards falsifiable interpretability research
M. L. Leavitt and A. Morcos · 2020
Cited alongside, same era.
interpreting gpt: the logit lens, 2020
nostalgebraist · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter · 2020
Cited alongside, same era.
Inducing causal structure for interpretable neural networks
A. Geiger, Z. Wu, H. Lu, J. Rozner, E. Kreiss, T. Icard, N. Goodman, and C. Potts · 2022
Later among the works it cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Later among the works it cites.
Causal abstraction with soft interventions
R. Massidda, A. Geiger, T. Icard, and D. Bacciu · 2022
Later among the works it cites.
Acquisition of chess knowledge in alphazero
T. McGrath, A. Kapishnikov, N. Tomašev, A. Pearce, M. Wattenberg, D. Hassabis, B. Kim, U. Paquet, and V. Kramnik · 2022
Later among the works it cites.
Locating and editing factual associations in gpt
K. Meng, D. Bau, A. Andonian, and Y. Belinkov · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2021
Cited alongside, same era.
Causal analysis of syntactic agreement mechanisms in neural language models
M. Finlayson, A. Mueller, S. Gehrmann, S. Shieber, T. Linzen, and Y. Belinkov · 2021
Cited alongside, same era.
Causal abstractions of neural networks
A. Geiger, H. Lu, T. Icard, and C. Potts · 2021
Cited alongside, same era.
Multimodal neurons in artificial neural networks
G. Goh, N. C. †, C. V. †, S. Carter, M. Petrov, L. Schubert, A. Radford, and C. Olah · 2021
Cited alongside, same era.
Causal scrubbing, a method for rigorously testing interpretability hypotheses
L. Chan, A. Garriga-Alonso, N. Goldwosky-Dill, R. Greenblatt, J. Nitishinskaya, A. Radhakrishnan, B. Shlegeris, and N. Thomas · 2022
Cited alongside, same era.
Analyzing transformers in embedding space
G. Dar, M. Geva, A. Gupta, and J. Berant · 2022
Cited alongside, same era.
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, et al · 2022
Later among the works it cites.
Formal algorithms for transformers
M. Phuong and M. Hutter · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
A. Power, Y. Burda, H. Edwards, I. Babuschkin, and V. Misra · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2022
Later among the works it cites.
Eliciting latent predictions from transformers with the tuned lens
N. Belrose, Z. Furman, L. Smith, D. Halawi, I. Ostrovsky, L. McKinney, S. Biderman, and J. Steinhardt · 2023
Closest in time.
A toy model of universality: Reverse engineering how networks learn group operations
B. Chughtai, L. Chan, and N. Nanda · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability
A. Conmy, A. N. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing
W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
N. Nanda, L. Chan, T. Liberum, J. Smith, and J. Steinhardt · 2023
Closest in time.