Fetching the paper…
Reading the bibliography…
Activation patching is a popular mechanistic interpretability technique, but has many subtleties regarding how it is applied and how one may interpret the results.
“Discovering the Compositional Structure of Vector Representations with Role Learning Networks”
Paul Soulos, Tom McCoy, Tal Linzen and Paul Smolensky · 1910
Earlier work this paper cites.
Atticus Geiger, Kyle Richardson and Christopher Potts · 2004
Earlier work this paper cites.
“Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias”
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer and Stuart Shieber · 2004
Earlier work this paper cites.
“Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models”
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen and Yonatan Belinkov · 2021
Earlier work this paper cites.
“Causal Abstractions of Neural Networks”
Atticus Geiger, Hanson Lu, Thomas Icard and Christopher Potts · 2021
Earlier work this paper cites.
“Inducing Causal Structure for Interpretable Neural Networks”
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah. Goodman and Christopher Potts · 2021
Earlier work this paper cites.
Peter Hase, Harry Xie and Mohit Bansal · 2021
Earlier work this paper cites.
“Causal scrubbing: A method for rigorously testing interpretability hypotheses”, Alignment Forum, 2022
Lawrence Chan, Adria Garriga-Alonso, Nix Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris and Nate Thomas · 2022
Earlier work this paper cites.
“Locating and Editing Factual Associations in GPT”
Kevin Meng, David Bau, Alex Andonian and Yonatan Belinkov · 2022
Earlier work this paper cites.
“Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small”
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris and Jacob Steinhardt · 2022
Earlier work this paper cites.
“Towards Automated Circuit Discovery for Mechanistic Interpretability”
Arthur Conmy, Augustine. Mavor-Parker, Aengus Lynch, Stefan Heimersheim and Adrià Garriga-Alonso · 2023
Cited alongside, same era.
“Sparse Autoencoders Find Highly Interpretable Features in Language Models”
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben and Lee Sharkey · 2023
Cited alongside, same era.
“How do Language Models Bind Entities in Context?”
Jiahai Feng and Jacob Steinhardt · 2023
Cited alongside, same era.
“Dissecting Recall of Factual Associations in Auto-Regressive Language Models”
Mor Geva, Jasmijn Bastings, Katja Filippova and Amir Globerson · 2023
Cited alongside, same era.
Tom Lieberum, Matthew Rahtz, János Kramár, Neel Nanda, Geoffrey Irving, Rohin Shah and Vladimir Mikulik · 2023
Later among the works it cites.
“Copy Suppression: Comprehensively Understanding an Attention Head”
Callum McDougall, Arthur Conmy, Cody Rushing, Thomas McGrath and Neel Nanda · 2023
Later among the works it cites.
“The Hydra Effect: Emergent Self-repair in Language Model Computations”
Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik and Shane Legg · 2023
Later among the works it cites.
“Circuit Component Reuse Across Tasks in Transformer Language Models”
Jack Merullo, Carsten Eickhoff and Ellie Pavlick · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato and Aryaman Arora · 2023
Cited alongside, same era.
Michael Hanna, Ollie Liu and Alexandre Variengien · 2023
Cited alongside, same era.
Peter Hase, Mohit Bansal, Been Kim and Asma Ghandeharioun · 2023
Cited alongside, same era.
“A circuit for Python docstrings in a 4-layer attention-only transformer”, 2023
Stefan Heimersheim and Jett Janiak · 2023
Cited alongside, same era.
“In-Context Learning Creates Task Vectors”
Roee Hendel, Mor Geva and Amir Globerson · 2023
Cited alongside, same era.
“Rigorously Assessing Natural Language Explanations of Neurons”
Jing Huang, Atticus Geiger, Karel D’Oosterlinck, Zhengxuan Wu and Christopher Potts · 2023
Cited alongside, same era.
“Attribution Patching: Activation Patching At Industrial Scale” Section “How to Think About Activation Patching”
Neel Nanda · 2023
Later among the works it cites.
“Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level (Post 1)”, Alignment Forum, 2023
Neel Nanda, SenR, János Kramár and Rohin Shah · 2023
Later among the works it cites.
Alessandro Stolfo, Yonatan Belinkov and Mrinmaya Sachan · 2023
Later among the works it cites.
“Linear Representations of Sentiment in Large Language Models”
Curt Tigges, Oskar Hollinsworth, Atticus Geiger and Neel Nanda · 2023
Later among the works it cites.
“Function Vectors in Large Language Models”
Eric Todd, Millicent. Li, Arnab Sharma, Aaron Mueller, Byron. Wallace and David Bau · 2023
Later among the works it cites.
“Towards Best Practices of Activation Patching in Language Models: Metrics and Methods”
Fred Zhang and Neel Nanda · 2023
Later among the works it cites.