Fetching the paper…
Reading the bibliography…
We show how to "compile" human-readable programs into standard decoder-only transformer models.
K-SVD: An algorithm for designing overcomplete dictionaries for sparse representation
M. Aharon, M. Elad, and A. Bruckstein · 2006
Earlier work this paper cites.
Compressed sensing
D. L. Donoho · 2006
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Z. C. Lipton · 2018
Earlier work this paper cites.
Analysis methods in neural language processing: A survey
Y. Belinkov and J. Glass · 2019
Earlier work this paper cites.
Machine learning interpretability: A survey on methods and metrics
D. V. Carvalho, E. M. Pereira, and J. S. Cardoso · 2019
Earlier work this paper cites.
Evaluating explanation without ground truth in interpretable machine learning
F. Yang, M. Du, and X. Hu · 2019
Earlier work this paper cites.
Benchmarking attribution methods with relative feature importance
M. Yang and B. Kim · 2019
Earlier work this paper cites.
Debugging tests for model explanations
J. Adebayo, M. Muelly, I. Liccardi, and B. Kim · 2020
Earlier work this paper cites.
Thread: Circuits
N. Cammarata, S. Carter, G. Goh, C. Olah, M. Petrov, L. Schubert, C. Voss, B. Egan, and S. K. Lim · 2020
Earlier work this paper cites.
A survey of the state of explainable AI for natural language processing
M. Danilevsky, K. Qian, R. Aharonov, Y. Katsis, B. Kawas, and P. Sen · 2020
Earlier work this paper cites.
Haiku: Sonnet for JAX, 2020
T. Hennigan, T. Cai, T. Norman, and I. Babuschkin · 2020
Earlier work this paper cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
A. Jacovi and Y. Goldberg · 2020
Cited alongside, same era.
Towards falsifiable interpretability research
M. L. Leavitt and A. Morcos · 2020
Cited alongside, same era.
A primer in BERTology: What we know about how BERT works
A. Rogers, O. Kovaleva, and A. Rumshisky · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2021
Cited alongside, same era.
A multidisciplinary survey and framework for design and evaluation of explainable AI systems
S. Mohseni, N. Zarei, and E. D. Ragan · 2021
Cited alongside, same era.
Thinking like transformers
Mechanistic interpretability, variables, and the importance of interpretable bases
C. Olah · 2022
Later among the works it cites.
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, S. Johnston, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2022
Later among the works it cites.
Polysemanticity and capacity in neural networks
A. Scherlis, K. Sachan, A. S. Jermyn, J. Benton, and B. Shlegeris · 2022
Later among the works it cites.
Statistically meaningful approximation: a case study on approximating Turing machines with transformers
C. Wei, Y. Chen, and T. Ma · 2022
Later among the works it cites.
Do feature attribution methods correctly attribute features?
Y. Zhou, S. Booth, M. T. Ribeiro, and J. Shah · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Weiss, Y. Goldberg, and E. Yahav · 2021
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Y. Belinkov · 2022
Cited alongside, same era.
Causal scrubbing, a method for rigorously testing interpretability hypotheses
L. Chan, A. Garriga-Alonso, N. Goldwosky-Dill, R. Greenblatt, J. Nitishinskaya, A. Radhakrishnan, B. Shlegeris, and N. Thomas · 2022
Cited alongside, same era.
Softmax linear units
N. Elhage, T. Hume, C. Olsson, N. Nanda, T. Henighan, S. Johnston, S. ElShowk, N. Joseph, N. DasSarma, B. Mann, D. Hernandez, A. Askell, K. Ndousse, A. Jones, D. Drain, A. Chen, Y. Bai, D. Ganguli, L. Lovitt, Z. Hatfield-Dodds, J. Kernion, T. Conerly, S. Kravec, S. Fort, S. Kadavath, J. Jacobson, E. Tran-Johnson, J. Kaplan, J. Clark, T. Brown, S. McCandlish, D. Amodei, and C. Olah · 2022
Cited alongside, same era.
Toy models of superposition
N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, R. Grosse, S. McCandlish, J. Kaplan, D. Amodei, M. Wattenberg, and C. Olah · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
K. Meng, D. Bau, A. J. Andonian, and Y. Belinkov · 2022
Cited alongside, same era.
Saturated transformers are constant-depth threshold circuits
W. Merrill, A. Sabharwal, and N. A. Smith · 2022
Cited alongside, same era.
What learning algorithm is in-context learning? Investigations with linear models
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability
A. Conmy, A. N. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso · 2023
Closest in time.
Learning transformer programs
D. Friedman, A. Wettig, and D. Chen · 2023
Closest in time.
Looped transformers as programmable computers
A. Giannou, S. Rajput, J.-y. Sohn, K. Lee, J. D. Lee, and D. Papailiopoulos · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
N. Nanda, L. Chan, T. Liberum, J. Smith, and J. Steinhardt · 2023
Closest in time.
Toward transparent AI: A survey on interpreting the inner structures of deep neural networks
T. Räukur, A. Ho, S. Casper, and D. Hadfield-Menell · 2023
Closest in time.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2023
Closest in time.