Fetching the paper…
Reading the bibliography…
Large language models (LLMs) demonstrate surprising capabilities, but we do not understand how they are implemented.
Multiple Comparisons among Means
O. J. Dunn · 1961
Earlier work this paper cites.
Hypothesis Testing , pages 65–80
G. A. Young and R. L. Smith · 2005
Earlier work this paper cites.
A Kernel Statistical Test of Independence
A. Gretton, K. Fukumizu, C. Teo, L. Song, B. Schölkopf, and A. Smola · 2007
Earlier work this paper cites.
Towards a Rigorous Science of Interpretable Machine Learning
F. Doshi-Velez and B. Kim · 2017
Earlier work this paper cites.
OpenWebText Corpus
A. Gokaslan and V. Cohen · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?
A. Jacovi and Y. Goldberg · 2020
Earlier work this paper cites.
Explainable AI: A Review of Machine Learning Interpretability Methods
P. Linardatos, V. Papastefanopoulos, and S. B. Kotsiantis · 2020
Earlier work this paper cites.
Compositional Explanations of Neurons
J. Mu and J. Andreas · 2020
Earlier work this paper cites.
Zoom In: An Introduction to Circuits
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter · 2020
Earlier work this paper cites.
High-Low Frequency Detectors
L. Schubert, C. Voss, N. Cammarata, G. Goh, and C. Olah · 2020
Earlier work this paper cites.
The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks
Z. Zhong, Z. Liu, M. Tegmark, and J. Andreas · 2020
Earlier work this paper cites.
Curve Circuits
N. Cammarata, G. Goh, S. Carter, C. Voss, L. Schubert, and C. Olah · 2021
Cited alongside, same era.
A Mathematical Framework for Transformer Circuits, 2021
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2021
Cited alongside, same era.
Causal Abstractions of Neural Networks
A. Geiger, H. Lu, T. Icard, and C. Potts · 2021
Cited alongside, same era.
Thinking Like Transformers, 2021
G. Weiss, Y. Goldberg, and E. Yahav · 2021
Cited alongside, same era.
Causal Scrubbing: A Method for Rigorously Testing Interpretability Hypotheses
L. Chan, A. Garriga-Alonso, N. Goldowsky-Dill, R. Greenblatt, J. Nitishinskaya, A. Radhakrishnan, B. Shlegeris, and N. Thomas · 2022
Cited alongside, same era.
TransformerLens
N. Nanda and J. Bloom · 2022
ALMANACS: A Simulatability Benchmark for Language Model Explainability
E. Mills, S. Su, S. Russell, and S. Emmons · 2023
Later among the works it cites.
Understanding Addition in Transformers
P. Quirke et al · 2023
Later among the works it cites.
FIND: A Function Description Benchmark for Evaluating Interpretability Methods
S. Schwettmann, T. R. Shaham, J. Materzynska, N. Chowdhury, S. Li, J. Andreas, D. Bau, and A. Torralba · 2023
Later among the works it cites.
Grokking Group Multiplication with Cosets
D. Stander, Q. Yu, H. Fan, and S. Biderman · 2023
Later among the works it cites.
Attribution Patching Outperforms Automated Circuit Discovery
A. Syed, C. Rager, and A. Conmy · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
In-Context Learning and Induction Heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, et al · 2022
Cited alongside, same era.
Red Teaming Deep Neural Networks with Feature Synthesis Tools
S. Casper, T. Bu, Y. Li, J. Li, K. Zhang, K. Hariharan, and D. Hadfield-Menell · 2023
Cited alongside, same era.
Towards Automated Circuit Discovery for Mechanistic Interpretability
A. Conmy, A. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso · 2023
Cited alongside, same era.
How Does GPT-2 Compute Greater-Than
M. Hanna, O. Liu, and A. Variengien · 2023
Cited alongside, same era.
A Circuit for Python Docstrings in a 4-layer Attention-Only Transformer
S. Heimersheim and J. Janiak · 2023
Cited alongside, same era.
T. Lieberum, M. Rahtz, J. Kramár, N. Nanda, G. Irving, R. Shah, and V. Mikulik · 2023
Cited alongside, same era.
Later among the works it cites.
Look Before You Leap: A Universal Emergent Decomposition of Retrieval Tasks in Language Models
A. Variengien and E. Winsor · 2023
Later among the works it cites.
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small
K. R. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2023
Later among the works it cites.
Learning Transformer Programs
D. Friedman, A. Wettig, and D. Chen · 2024
Closest in time.
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
A. Geiger, Z. Wu, C. Potts, T. Icard, and N. Goodman · 2024
Closest in time.
Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models
P. Hase, M. Bansal, B. Kim, and A. Ghandeharioun · 2024
Closest in time.
Tracr: Compiled Transformers as a Laboratory for Interpretability
D. Lindner, J. Kramár, S. Farquhar, M. Rahtz, T. McGrath, and V. Mikulik · 2024
Closest in time.
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
F. Zhang and N. Nanda · 2024
Closest in time.