Fetching the paper…
Reading the bibliography…
\emph{Circuit analysis} is a promising technique for understanding the internal mechanisms of language models.
Causal diagrams for empirical research
J. Pearl · 1995
Earlier work this paper cites.
J. Pearl · 2012
Earlier work this paper cites.
Feature visualization
C. Olah, A. Mordvintsev, and L. Schubert · 2017
Earlier work this paper cites.
G. Irving, P. Christiano, and D. Amodei · 2018
Earlier work this paper cites.
The building blocks of interpretability
C. Olah, A. Satyanarayan, I. Johnson, S. Carter, L. Schubert, K. Ye, and A. Mordvintsev · 2018
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov · 2019
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems
E. Hubinger, C. van Merwijk, V. Mikulik, J. Skalse, and S. Garrabrant · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
P. Michel, O. Levy, and G. Neubig · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Thread: Circuits
N. Cammarata, S. Carter, G. Goh, C. Olah, M. Petrov, L. Schubert, C. Voss, B. Egan, and S. K. Lim · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Understanding RL vision
J. Hilton, N. Cammarata, S. Carter, G. Goh, and C. Olah · 2020
Earlier work this paper cites.
interpreting GPT: the logit lens
nostalgebraist · 2020
Earlier work this paper cites.
Low-complexity probing via finding subnetworks
S. Cao, V. Sanh, and A. M. Rush · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
A. Geiger, H. Lu, T. Icard, and C. Potts · 2021
Earlier work this paper cites.
LoRa: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Z. Kenton, T. Everitt, L. Weidinger, I. Gabriel, V. Mikulik, and G. Irving · 2021
Cited alongside, same era.
Constitutional AI: Harmlessness from AI feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al · 2022
Cited alongside, same era.
Causal scrubbing: a method for rigorously testing interpretability hypotheses
L. Chan, A. Garriga-Alonso, N. Goldowsky-Dill, R. Greenblatt, J. Nitishinskaya, A. Radhakrishnan, B. Shlegeris, and N. Thomas · 2022
Cited alongside, same era.
Analyzing transformers in embedding space
G. Dar, M. Geva, A. Gupta, and J. Berant · 2022
Cited alongside, same era.
Softmax linear units
N. Elhage, T. Hume, C. Olsson, N. Nanda, T. Henighan, S. Johnston, S. ElShowk, N. Joseph, N. DasSarma, B. Mann, D. Hernandez, A. Askell, K. Ndousse, A. Jones, D. Drain, A. Chen, Y. Bai, D. Ganguli, L. Lovitt, Z. Hatfield-Dodds, J. Kernion, T. Conerly, S. Kravec, S. Fort, S. Kadavath, J. Jacobson, E. Tran-Johnson, J. Kaplan, J. Clark, T. Brown, S. McCandlish, D. Amodei, and C. Olah · 2022
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2022
Later among the works it cites.
Eliciting latent predictions from transformers with the tuned lens
N. Belrose, Z. Furman, L. Smith, D. Halawi, L. McKinney, I. Ostrovsky, S. Biderman, and J. Steinhardt · 2023
Closest in time.
Language models can explain neurons in language models
S. Bills, N. Cammarata, D. Mossing, H. Tillman, L. Gao, G. Goh, I. Sutskever, J. Leike, J. Wu, and W. Saunders · 2023
Closest in time.
Decision transformer interpretability
J. I. Bloom and P. Colognese · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability
A. Conmy, A. N. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Toy models of superposition
N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, R. Grosse, S. McCandlish, J. Kaplan, D. Amodei, M. Wattenberg, and C. Olah · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
M. Geva, A. Caciularu, K. R. Wang, and Y. Goldberg · 2022
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgements
A. Glaese, N. McAleese, M. Trębacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, et al · 2022
Cited alongside, same era.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Cited alongside, same era.
Language model compression with weighted low-rank factorization
Y.-C. Hsu, T. Hua, S. Chang, Q. Lou, Y. Shen, and H. Jin · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
K. Meng, D. Bau, A. Andonian, and Y. Belinkov · 2022
Cited alongside, same era.
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, S. Johnston, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2022
Cited alongside, same era.
Closest in time.
Jump to conclusions: Short-cutting transformers with linear transformations
A. Y. Din, T. Karidi, L. Choshen, and M. Geva · 2023
Closest in time.
Dissecting recall of factual associations in auto-regressive language models
M. Geva, J. Bastings, K. Filippova, and A. Globerson · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing
W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas · 2023
Closest in time.
A circuit for Python docstrings in a 4-layer attention-only transformer
S. Heimersheim and J. Janiak · 2023
Closest in time.
A comparison of causal scrubbing, causal abstractions, and related methods
E. Jenner, A. Garriga-Alonso, and E. Zverev · 2023
Closest in time.
Circuits updates — May 2023: Attention head superposition
A. Jermyn, C. Olah, and T. Henighan · 2023
Closest in time.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
N. Nanda, L. Chan, T. Liberum, J. Smith, and J. Steinhardt · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Model evaluation for extreme risks
T. Shevlane, S. Farquhar, B. Garfinkel, M. Phuong, J. Whittlestone, J. Leung, D. Kokotajlo, N. Marchal, M. Anderljung, N. Kolt, et al · 2023
Closest in time.
Interpretability at scale: Identifying causal mechanisms in alpaca
Z. Wu, A. Geiger, C. Potts, and N. D. Goodman · 2023
Closest in time.