Fetching the paper…
Reading the bibliography…
Mechanistic interpretability aims to understand the behavior of neural networks by reverse-engineering their internal computations.
Learning factorial codes by predictability minimization
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Stochastic estimation with z2 noise
Shao-Jing Dong and Keh-Fei Liu · 1994
Earlier work this paper cites.
Paths and consistency in additive cost sharing
Eric Friedman · 2004
Earlier work this paper cites.
Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer, and Stuart Shieber · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Disentangling factors of variation via generative entangling
Guillaume Desjardins, Aaron Courville, and Yoshua Bengio · 2012
Earlier work this paper cites.
Transformer Feed-Forward Layers Are Key-Value Memories, September 2021
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives, 2014
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2014
Earlier work this paper cites.
How far can we go without convolution: Improving fully-connected networks
Zhouhan Lin, Roland Memisevic, and Kishore Konda · 2015
Earlier work this paper cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Lattice Quantum Chromodynamics: Practical Essentials
F. Knechtli, M. Günther, and M. Peardon · 2016
Earlier work this paper cites.
Anh Nguyen, Jason Yosinski, and Jeff Clune · 2016
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks, 2017
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Disentangling by factorising
Hyunjik Kim and Andriy Mnih · 2018
Cited alongside, same era.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
The hessian penalty: A weak prior for unsupervised disentanglement
William Peebles, John Peebles, Jun-Yan Zhu, Alexei Efros, and Antonio Torralba · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Later among the works it cites.
A toy model of universality: Reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Later among the works it cites.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Later among the works it cites.
Tinystories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Cited alongside, same era.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Cited alongside, same era.
Explaining neural networks by decoding layer activations
Johannes Schneider and Michalis Vlachos · 2021
Cited alongside, same era.
Causal scrubbing: A method for rigorously testing interpretability hypotheses
Lawrence Chan, Adria Garriga-Alonso, Nix Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas · 2022
Cited alongside, same era.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Cited alongside, same era.
The singular value decompositions of transformer weight matrices are highly interpretable, Nov 2022
Beren Millidge and Sid Black · 2022
Cited alongside, same era.
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora · 2023
Later among the works it cites.
Finding neurons in a haystack: Case studies with sparse probing, 2023
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas · 2023
Later among the works it cites.
Polysemantic attention head in a 4-layer transformer, Nov 2023
Jett Janiak, Chris Mathwin, and Stefan Heimersheim · 2023
Later among the works it cites.
Locating and editing factual associations in gpt, 2023
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2023
Later among the works it cites.
Gpt-2’s positional embedding matrix is a helix, Jul 2023
Adam Yedidia · 2023
Later among the works it cites.
The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2023
Later among the works it cites.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
SAE Visualizer
Callum McDougall · 2024
Closest in time.
Toward a mathematical framework for computation in superposition, Jan 2024
Dmitry Vaintrob, Jake Mendel, and Kaarel Hänni · 2024
Closest in time.
From louvain to leiden: guaranteeing well-connected communities
V. A. Traag, L. Waltman, and N. J. van Eck · 2045
Closest in time.