Fetching the paper…
Reading the bibliography…
Interpretability research takes counterfactual theories of causality for granted.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 1910
Earlier work this paper cites.
Causes and conditions
J. L. Mackie · 1965
Earlier work this paper cites.
Structural equation methods in the social sciences
Arthur S. Goldberger · 1972
Earlier work this paper cites.
Counterfactuals
David K. Lewis · 1973
Earlier work this paper cites.
Scientific Explanation and the Causal Structure of the World
Wesley C. Salmon · 1984
Earlier work this paper cites.
Causation
David Lewis · 1986
Earlier work this paper cites.
Identifiability and exchangeability for direct and indirect effects
James M. Robins and Sander Greenland · 1992
Earlier work this paper cites.
Causality without counterfactuals
Wesley C. Salmon · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Causation as influence
David Lewis · 2000
Earlier work this paper cites.
Causality: Models, Reasoning, and Inference
Judea Pearl · 2000
Earlier work this paper cites.
The intransitivity of causation revealed in equations and graphs
Christopher Hitchcock · 2001
Earlier work this paper cites.
Direct and indirect effects
Judea Pearl · 2001
Earlier work this paper cites.
Two concepts of causation
Ned Hall · 2004
Earlier work this paper cites.
Causal inference using the algorithmic markov condition
Dominik Janzing and Bernhard Schölkopf · 2010
Earlier work this paper cites.
Causation: A User’s Guide
L. A. Paul and Ned Hall · 2013
Earlier work this paper cites.
Sufficient conditions for causality to be transitive
Joseph Y Halpern · 2016
Earlier work this paper cites.
The transitivity and asymmetry of actual causation
Sander Beckers and Joost Vennekens · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema · 2018
Earlier work this paper cites.
Analyzing redundancy in pretrained transformer models
Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov · 2020
Earlier work this paper cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Earlier work this paper cites.
Discovering the compositional structure of vector representations with role learning networks
Paul Soulos, R. Thomas McCoy, Tal Linzen, and Paul Smolensky · 2020
Cited alongside, same era.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Cited alongside, same era.
A Ramsey test analysis of causation for causal models
Holger Andreas and Mario Günther · 2021
Cited alongside, same era.
Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Michael Hanna, Ollie Liu, and Alexandre Variengien · 2023
Later among the works it cites.
Does localization inform editing? Surprising differences in causality-based localization vs. knowledge editing in language models
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun · 2023
Later among the works it cites.
Rigorously assessing natural language explanations of neurons
Jing Huang, Atticus Geiger, Karel D’Oosterlinck, Zhengxuan Wu, and Christopher Potts · 2023
Later among the works it cites.
Circuit breaking: Removing model behaviors with targeted ablation
Maximilian Li, Xander Davies, and Max Nadeau · 2023
Later among the works it cites.
Tom Lieberum, Matthew Rahtz, János Kramár, Neel Nanda, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Causal analysis of syntactic agreement mechanisms in neural language models
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov · 2021
Cited alongside, same era.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas F Icard, and Christopher Potts · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Cited alongside, same era.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Kamal Saab, Tri Dao, Atri Rudra, and Christopher Re · 2021
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg · 2022
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Re · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
Later among the works it cites.
The hydra effect: Emergent self-repair in language model computations
Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg · 2023
Later among the works it cites.
Mass editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau · 2023
Later among the works it cites.
A comprehensive mechanistic interpretability explainer & glossary
Neel Nanda · 2023
Later among the works it cites.
Attribution patching outperforms automated circuit discovery
Aaquib Syed, Can Rager, and Arthur Conmy · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Later among the works it cites.
Interpretability at scale: Identifying causal mechanisms in alpaca
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah Goodman · 2023
Later among the works it cites.
Characterizing mechanisms for factual recall in language models
Qinan Yu, Jack Merullo, and Ellie Pavlick · 2023
Later among the works it cites.
A regularity theory of causation
Holger Andreas and Mario Günther · 2024
Closest in time.
Tweet in response to Alexander Rush on the rise of mechanistic interpretability, 2024
Jacob Andreas · 2024
Closest in time.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2024
Closest in time.
Information flow routes: Automatically interpreting language models at scale
Javier Ferrando and Elena Voita · 2024
Closest in time.
Finding alignments between interpretable causal variables and distributed neural representations
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D. Goodman · 2024
Closest in time.
Successor heads: Recurring, interpretable attention heads in the wild
Rhys Gould, Euan Ong, George Ogden, and Arthur Conmy · 2024
Closest in time.
How to use and interpret activation patching
Stefan Heimersheim and Neel Nanda · 2024
Closest in time.
Atp*: An efficient and scalable method for localizing llm behaviour to components
János Kramár, Tom Lieberum, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Sparsify: A mechanistic interpretability research agenda, 2024
Lee Sharkey · 2024
Closest in time.
Function vectors in large language models
Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau · 2024
Closest in time.