Fetching the paper…
Reading the bibliography…
Mechanistic interpretability work attempts to reverse engineer the learned algorithms present inside neural networks.
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi · 2018
Earlier work this paper cites.
Analyzing and interpreting neural networks for nlp: A report on the first blackboxnlp workshop
Afra Alishahi, Grzegorz Chrupała, and Tal Linzen · 2019
Earlier work this paper cites.
Curve circuits
Nick Cammarata, Gabriel Goh, Shan Carter, Chelsea Voss, Ludwig Schubert, and Chris Olah · 2020
Earlier work this paper cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Causal scrubbing: a method for rigorously testing interpretability hypotheses [redwood research], 2022
Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
Transformerlens
Neel Nanda and Joseph Bloom · 2022
Earlier work this paper cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Earlier work this paper cites.
Towards best practices of activation patching in language models: Metrics and methods, 2024
Fred Zhang and Neel Nanda · 2022
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling, 2023
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal · 2023
Earlier work this paper cites.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Cited alongside, same era.
A toy model of universality: Reverse engineering how networks learn group operations, 2023
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models, 2023
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Attribution patching outperforms automated circuit discovery, 9 2023
Aaquib Syed and Can Rager · 2023
Later among the works it cites.
Linear representations of sentiment in large language models, 2023
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization, 2023
Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Later among the works it cites.
Chirag Agarwal, Sree Harsha Tanneru, and Himabindu Lakkaraju · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Finding alignments between interpretable causal variables and distributed neural representations
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas F. Icard, and Noah D. Goodman · 2023
Cited alongside, same era.
Localizing model behavior with path patching, 2023
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora · 2023
Cited alongside, same era.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model, 2023
Michael Hanna, Ollie Liu, and Alexandre Variengien · 2023
Cited alongside, same era.
A circuit for Python docstrings in a 4-layer attention-only transformer, 2023
Stefan Heimersheim and Jett Janiak · 2023
Cited alongside, same era.
Emergent world representations: Exploring a sequence model trained on a synthetic task, 2023
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Cited alongside, same era.
Tracr: Compiled transformers as a laboratory for interpretability
David Lindner, János Kramar, Matthew Rahtz, Thomas McGrath, and Vladimir Mikulik · 2023
Cited alongside, same era.
Is this the subspace you are looking for? an interpretability illusion for subspace activation patching, 2023
Aleksandar Makelov, Georg Lange, and Neel Nanda · 2023
Cited alongside, same era.
Language models represent space and time, 2024
Wes Gurnee and Max Tegmark · 2024
Closest in time.
Atp*: An efficient and scalable method for localizing llm behaviour to components, 2024
Janos Kramar, Tom Lieberum, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Circuit breaking: Removing model behaviors with targeted ablation, 2024
Maximilian Li, Xander Davies, and Max Nadeau · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Steering llama 2 via contrastive activation addition, 2024
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner · 2024
Closest in time.
Sparsify: A mechanistic interpretability research agenda, 2024
Lee Sharkey · 2024
Closest in time.
The transformer model in equations, 2024
John Thickstun · 2024
Closest in time.