Fetching the paper…
Reading the bibliography…
Activation Patching is a method of directly computing causal attributions of behavior to model components.
The generalization of ‘Student’s’ problem when several different population variances are involved
B. L. Welch · 1947
Earlier work this paper cites.
Identifiability and exchangeability for direct and indirect effects
J. M. Robins and S. Greenland · 1992
Earlier work this paper cites.
Causality: Models, Reasoning and Inference
J. Pearl · 2000
Earlier work this paper cites.
Direct and indirect effects, 2001
J. Pearl · 2001
Earlier work this paper cites.
Perforatedcnns: Acceleration through elimination of redundant convolutions
M. Figurnov, A. Ibraimova, D. P. Vetrov, and P. Kohli · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
A. Veit, M. J. Wilber, and S. Belongie · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
P. Molchanov, S. Tyree, T. Karras, T. Aila, and J. Kautz · 2017
Earlier work this paper cites.
Attention is all you need, 2017
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Learning sparse neural networks through l 0 l_{0} regularization, 2018
C. Louizos, M. Welling, and D. P. Kingma · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
Are sixteen heads really better than one?
P. Michel, O. Levy, and G. Neubig · 2019
Earlier work this paper cites.
Thread: Circuits
N. Cammarata, S. Carter, G. Goh, C. Olah, M. Petrov, L. Schubert, C. Voss, B. Egan, and S. K. Lim · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy · 2020
Earlier work this paper cites.
Neural natural language inference models partially embed theories of lexical entailment and negation, 2020
A. Geiger, K. Richardson, and C. Potts · 2020
Earlier work this paper cites.
interpreting gpt: the logit lens
nostalgebraist · 2020
Earlier work this paper cites.
Discovering the compositional structure of vector representations with role learning networks
P. Soulos, R. T. McCoy, T. Linzen, and P. Smolensky · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
J. Vig, S. Gehrmann, Y. Belinkov, S. Qian, D. Nevo, Y. Singer, and S. Shieber · 2020
Earlier work this paper cites.
Sparse interventions in language models with differentiable masking, 2021
N. D. Cao, L. Schmid, D. Hupkes, and I. Titov · 2021
Earlier work this paper cites.
Causal analysis of syntactic agreement mechanisms in neural language models
M. Finlayson, A. Mueller, S. Gehrmann, S. Shieber, T. Linzen, and Y. Belinkov · 2021
Earlier work this paper cites.
Causal abstractions of neural networks, 2021
A. Geiger, H. Lu, T. Icard, and C. Potts · 2021
Cited alongside, same era.
Causal scrubbing, a method for rigorously testing interpretability hypotheses
L. Chan, A. Garriga-Alonso, N. Goldwosky-Dill, R. Greenblatt, J. Nitishinskaya, A. Radhakrishnan, B. Shlegeris, and N. Thomas · 2022
Cited alongside, same era.
Inducing causal structure for interpretable neural networks, 2022
A. Geiger, Z. Wu, H. Lu, J. Rozner, E. Kreiss, T. Icard, N. D. Goodman, and C. Potts · 2022
Cited alongside, same era.
An empirical analysis of compute-optimal large language model training
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, O. Vinyals, J. Rae, and L. Sifre · 2022
Cited alongside, same era.
A comparison of causal scrubbing, causal abstractions, and related methods
E. Jenner, A. Garriga-Alonso, and E. Zverev · 2022
Cited alongside, same era.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models, 2023
P. Hase, M. Bansal, B. Kim, and A. Ghandeharioun · 2023
Later among the works it cites.
In-context learning creates task vectors, 2023
R. Hendel, M. Geva, and A. Globerson · 2023
Later among the works it cites.
Rigorously assessing natural language explanations of neurons, 2023
J. Huang, A. Geiger, K. D’Oosterlinck, Z. Wu, and C. Potts · 2023
Later among the works it cites.
Improving activation steering in language models with mean-centring, 2023
O. Jorgensen, D. Cope, N. Schoots, and M. Shanahan · 2023
Later among the works it cites.
Inference-time intervention: Eliciting truthful answers from a language model, 2023
K. Li, O. Patel, F. Viégas, H. Pfister, and M. Wattenberg · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attribution patching: Activation patching at industrial scale
N. Nanda · 2022
Cited alongside, same era.
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, S. Johnston, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2022
Cited alongside, same era.
Leace: Perfect linear concept erasure in closed form
N. Belrose, D. Schneider-Joseph, S. Ravfogel, R. Cotterell, E. Raff, and S. Biderman · 2023
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, A. Skowron, L. Sutawika, and O. van der Wal · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, Z. Hatfield-Dodds, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability, 2023
A. Conmy, A. N. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso · 2023
Cited alongside, same era.
T. Lieberum, M. Rahtz, J. Kramár, N. Nanda, G. Irving, R. Shah, and V. Mikulik · 2023
Later among the works it cites.
Copy suppression: Comprehensively understanding an attention head, 2023
C. McDougall, A. Conmy, C. Rushing, T. McGrath, and N. Nanda · 2023
Later among the works it cites.
The hydra effect: Emergent self-repair in language model computations, 2023
T. McGrath, M. Rahtz, J. Kramár, V. Mikulik, and S. Legg · 2023
Later among the works it cites.
Locating and editing factual associations in gpt, 2023
K. Meng, D. Bau, A. Andonian, and Y. Belinkov · 2023
Later among the works it cites.
Circuit component reuse across tasks in transformer language models, 2023
J. Merullo, C. Eickhoff, and E. Pavlick · 2023
Later among the works it cites.
Fact finding: Attempting to reverse-engineer factual recall on the neuron level, Dec 2023
N. Nanda, S. Rajamanoharan, J. Kramár, and R. Shah · 2023
Later among the works it cites.
Steering llama 2 via contrastive activation addition, 2023
N. Rimsky, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. M. Turner · 2023
Later among the works it cites.
A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis, 2023
A. Stolfo, Y. Belinkov, and M. Sachan · 2023
Later among the works it cites.
Attribution patching outperforms automated circuit discovery, 2023
A. Syed, C. Rager, and A. Conmy · 2023
Later among the works it cites.
Linear representations of sentiment in large language models, 2023
C. Tigges, O. J. Hollinsworth, A. Geiger, and N. Nanda · 2023
Later among the works it cites.
Function vectors in large language models, 2023
E. Todd, M. L. Li, A. S. Sharma, A. Mueller, B. C. Wallace, and D. Bau · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization, 2023
A. M. Turner, L. Thiergart, D. Udell, G. Leech, U. Mini, and M. MacDiarmid · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency, 2023
A. Zou, L. Phan, S. Chen, J. Campbell, P. Guo, R. Ren, A. Pan, X. Yin, M. Mazeika, A.-K. Dombrowski, S. Goel, N. Li, M. J. Byun, Z. Wang, A. Mallen, S. Basart, S. Koyejo, D. Song, M. Fredrikson, J. Z. Kolter, and D. Hendrycks · 2023
Later among the works it cites.