Fetching the paper…
Reading the bibliography…
Decomposing model activations into interpretable components is a key open problem in mechanistic interpretability.
The use of confidence or fiducial limits illustrated in the case of the binomial
C. J. Clopper and E. S. Pearson · 1934
Earlier work this paper cites.
The cost of using exact confidence intervals for a binomial proportion
M. Thulin · 1935
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
B. A. Olshausen and D. J. Field · 1997
Earlier work this paper cites.
K-svd: An algorithm for designing overcomplete dictionaries for sparse representation
M. Aharon, M. Elad, and A. Bruckstein · 2006
Earlier work this paper cites.
Compositional explanations of neurons
J. Mu and J. Andreas · 2006
Earlier work this paper cites.
Sparse and Redundant Representations: From Theory to Applications in Signal and Image Processing
M. Elad · 2010
Earlier work this paper cites.
Learning fast approximations of sparse coding
K. Gregor and Y. LeCun · 2010
Earlier work this paper cites.
Sparse autoencoder
A. Ng · 2011
Earlier work this paper cites.
Feature visualization
C. Olah, A. Mordvintsev, and L. Schubert · 2017
Earlier work this paper cites.
Spine: Sparse interpretable neural embeddings, 2017
A. Subramanian, D. Pruthi, H. Jhamtani, T. Berg-Kirkpatrick, and E. Hovy · 2017
Earlier work this paper cites.
Activation atlas
S. Carter, Z. Armstrong, L. Schubert, I. Johnson, and C. Olah · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned, 2019
E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov · 2019
Earlier work this paper cites.
Zoom in: An introduction to circuits
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter · 2020
Earlier work this paper cites.
Causal mediation analysis for interpreting neural nlp: The case of gender bias, 2020
J. Vig, S. Gehrmann, Y. Belinkov, S. Qian, D. Nevo, S. Sakenis, J. Huang, Y. Singer, and S. Shieber · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2021
Earlier work this paper cites.
Multimodal neurons in artificial neural networks
G. Goh, N. Cammarata, C. Voss, S. Carter, M. Petrov, L. Schubert, A. Radford, and C. Olah · 2021
Earlier work this paper cites.
N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, et al · 2022
Earlier work this paper cites.
Natural language descriptions of deep visual features, 2022
E. Hernandez, S. Schwettmann, D. Bau, T. Bagashvili, A. Torralba, and J. Andreas · 2022
Earlier work this paper cites.
TransformerLens
N. Nanda and J. Bloom · 2022
Earlier work this paper cites.
Mechanistic interpretability, variables, and the importance of interpretable bases
C. Olah · 2022
Cited alongside, same era.
In-context learning and induction heads, 2022
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, et al · 2022
Cited alongside, same era.
[interim research report] taking features out of superposition with sparse autoencoders
L. Sharkey, D. Braun, and B. Millidge · 2022
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, Z. Hatfield-Dodds, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah · 2023
Cited alongside, same era.
A toy model of universality: Reverse engineering how networks learn group operations, 2023
B. Chughtai, L. Chan, and N. Nanda · 2023
Cited alongside, same era.
Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors, 2023
Z. Yun, Y. Chen, B. A. Olshausen, and Y. LeCun · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency, 2023
A. Zou, L. Phan, S. Chen, J. Campbell, P. Guo, R. Ren, A. Pan, X. Yin, M. Mazeika, A.-K. Dombrowski, S. Goel, N. Li, M. J. Byun, Z. Wang, A. Mallen, S. Basart, S. Koyejo, D. Song, M. Fredrikson, J. Z. Kolter, and D. Hendrycks · 2023
Later among the works it cites.
Open Source Sparse Autoencoders for all Residual Stream Layers of GPT-2 Small
J. Bloom · 2024
Closest in time.
Summing up the facts: Additive mechanisms behind factual recall in llms, 2024
B. Chughtai, A. Cooney, and N. Nanda · 2024
Closest in time.
Interpreting clip’s image representation via text-based decomposition, 2024
Y. Gandelsman, A. A. Efros, and J. Steinhardt · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards automated circuit discovery for mechanistic interpretability, 2023
A. Conmy, A. N. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models, 2023
H. Cunningham, A. Ewart, L. Riggs, R. Huben, and L. Sharkey · 2023
Cited alongside, same era.
Localizing model behavior with path patching, 2023
N. Goldowsky-Dill, C. MacLeod, L. Sato, and A. Arora · 2023
Cited alongside, same era.
Successor heads: Recurring, interpretable attention heads in the wild, 2023
R. Gould, E. Ong, G. Ogden, and A. Conmy · 2023
Cited alongside, same era.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model, 2023
M. Hanna, O. Liu, and A. Variengien · 2023
Cited alongside, same era.
Polysemantic attention head in a 4-layer transformer
J. Janiak, C. Mathwin, and S. Heimersheim · 2023
Cited alongside, same era.
Attention head superposition
A. Jermyn, C. Olah, and T. Henighan · 2023
Cited alongside, same era.
Automatically identifying local and global circuits with linear computation graphs, 2024
X. Ge, F. Zhu, W. Shu, J. Wang, Z. He, and X. Qiu · 2024
Closest in time.
Gemma, 2024
Gemma Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, L. Sifre, M. Rivière, M. S. Kale, J. Love, P. Tafti, L. Hussenot, and et al · 2024
Closest in time.
Dictionary learning improves patch-free circuit discovery in mechanistic interpretability: A case study on othello-gpt, 2024
Z. He, X. Ge, Q. Tang, T. Sun, Q. Cheng, and X. Qiu · 2024
Closest in time.
How to use and interpret activation patching, 2024
S. Heimersheim and N. Nanda · 2024
Closest in time.
Sparse autoencoders work on attention layer outputs
C. Kissane, R. Krzyzanowski, A. Conmy, and N. Nanda · 2024
Closest in time.
Towards principled evaluations of sparse autoencoders for interpretability and control
A. Makelov, G. Lange, and N. Nanda · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models, 2024
S. Marks, C. Rager, E. J. Michaud, Y. Belinkov, D. Bau, and A. Mueller · 2024
Closest in time.
SAE Visualizer
C. McDougall · 2024
Closest in time.
The quantization model of neural scaling, 2024
E. J. Michaud, Z. Liu, U. Girit, and M. Tegmark · 2024
Closest in time.
[Summary] Progress Update #1 from the GDM Mech Interp Team
N. Nanda, A. Conmy, L. Smith, S. Rajamanoharan, T. Lieberum, J. Kramár, and V. Varma · 2024
Closest in time.
Circuits Updates - April 2024
C. Olah, S. Carter, A. Jermyn, J. Batson, T. Henighan, J. Lindsey, T. Conerly, A. Templeton, J. Marcus, T. Bricken, E. Ameisen, H. Cunningham, and A. Golubeva · 2024
Closest in time.
Improving dictionary learning with gated sparse autoencoders, 2024
S. Rajamanoharan, A. Conmy, L. Smith, T. Lieberum, V. Varma, J. Kramár, R. Shah, and N. Nanda · 2024
Closest in time.
Steering llama 2 via contrastive activation addition, 2024
N. Rimsky, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. M. Turner · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
A. Templeton, T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones, H. Cunningham, N. L. Turner, C. McDougall, M. MacDiarmid, C. D. Freeman, T. R. Sumers, E. Rees, J. Batson, A. Jermyn, S. Carter, C. Olah, and T. Henighan · 2024
Closest in time.