Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Later among the works it cites.
Implicit representations of meaning in neural language models
Belinda Z. Li, Maxwell I. Nye, and Jacob Andreas · 2021
Later among the works it cites.
Compositional abstraction error and a category of causal models
Eigil F. Rischel and Sebastian Weichwald · 2021
Later among the works it cites.
CEBaB: Estimating the causal effects of real-world concepts on NLP model behavior
Original
Eldar David Abraham, Karel D’Oosterlinck, Amir Feder, Yair Ori Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu · 2022
Later among the works it cites.
Causal scrubbing: a method for rigorously testing interpretability hypotheses, 2022
Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas · 2022
Later among the works it cites.
Inducing character-level structure in subword-based language models with Type-level Interchange Intervention Training
Original
Jing Huang, Zhengxuan Wu, Kyle Mahowald, and Christopher Potts · 2022
Later among the works it cites.
Towards faithful model explanation in NLP: A survey
Original
Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch · 2022
Later among the works it cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Later among the works it cites.
Linear adversarial concept erasure
Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan D Cotterell · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Original
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Later among the works it cites.
Causal Proxy Models for concept-based model explanations
Original
Zhengxuan Wu, Karel D’Oosterlinck, Atticus Geiger, Amir Zur, and Christopher Potts · 2022
Later among the works it cites.
Causal abstraction for faithful interpretation of AI models
Original
Atticus Geiger, Chris Potts, and Thomas Icard · 2023
Closest in time.
Causal abstraction with soft interventions
Riccardo Massidda, Atticus Geiger, Thomas Icard, and Davide Bacciu · 2023
Closest in time.