Fetching the paper…
Reading the bibliography…
A key goal in mechanistic interpretability is circuit analysis: finding sparse subgraphs of models corresponding to specific behaviors or capabilities.
Correlating neural and symbolic representations of language
Chrupała, G. and Alishahi, A · 1905
Earlier work this paper cites.
Language Models are Few-Shot Learners, July 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
The mythos of model interpretability, 2017
Lipton, Z. C · 2017
Earlier work this paper cites.
Feature visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Earlier work this paper cites.
OpenWebText Corpus, 2019
Gokaslan, A. and Cohen, V · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Zoom In: An Introduction to Circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
Investigating Gender Bias in Language Models Using Causal Mediation Analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S · 2020
Earlier work this paper cites.
An interpretability illusion for bert
Bolukbasi, T., Pearce, A., Yuan, A., Coenen, A., Reif, E., Viégas, F., and Wattenberg, M · 2021
Earlier work this paper cites.
A Mathematical Framework for Transformer Circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Earlier work this paper cites.
Causal Abstractions of Neural Networks, October 2021
Geiger, A., Lu, H., Icard, T., and Potts, C · 2021
Earlier work this paper cites.
Mechanistic Interpretability, Variables, and the Importance of Interpretable Bases, 2022
Chris Olah · 2022
Earlier work this paper cites.
Softmax linear units
Elhage, N., Hume, T., Olsson, C., Nanda, N., Henighan, T., Johnston, S., ElShowk, S., Joseph, N., DasSarma, N., Mann, B., Hernandez, D., Askell, A., Ndousse, K., Jones, A., Drain, D., Chen, A., Bai, Y., Ganguli, D., Lovitt, L., Hatfield-Dodds, Z., Kernion, J., Conerly, T., Kravec, S., Fort, S., Kadavath, S., Jacobson, J., Tran-Johnson, E., Kaplan, J., Clark, J., Brown, T., McCandlish, S., Amodei, D., and Olah, C · 2022
Earlier work this paper cites.
TransformerLens, 2022
Nanda, N. and Bloom, J · 2022
Earlier work this paper cites.
In-context Learning and Induction Heads, September 2022
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Earlier work this paper cites.
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Earlier work this paper cites.
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling, May 2023
Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., and van der Wal, O · 2023
Cited alongside, same era.
Language models can explain neurons in language models, 2023
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Cited alongside, same era.
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Cited alongside, same era.
Towards Automated Circuit Discovery for Mechanistic Interpretability, October 2023
Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Cited alongside, same era.
Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors, 2023
Yun, Z., Chen, Y., Olshausen, B. A., and LeCun, Y · 2023
Later among the works it cites.
Using features for easy circuit identification, 2024
Batson, J., Chen, B., and Jones, A · 2024
Closest in time.
Case Studies in Reverse-Engineering Sparse Autoencoder Features by Using MLP Linearization, 2024
Dunefsky, J., Chlenski, P., Rajamanoharan, S., and Nanda, N · 2024
Closest in time.
A Primer on the Inner Workings of Transformer-based Language Models, May 2024
Ferrando, J., Sarti, G., Bisazza, A., and Costa-jussà, M. R · 2024
Closest in time.
Interpreting CLIP’s Image Representation via Text-Based Decomposition, March 2024
Gandelsman, Y., Efros, A. A., and Steinhardt, J · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Cited alongside, same era.
Dunefsky, J. and Cohan, A · 2023
Cited alongside, same era.
Localizing Model Behavior with Path Patching, May 2023
Goldowsky-Dill, N., MacLeod, C., Sato, L., and Arora, A · 2023
Cited alongside, same era.
Successor Heads: Recurring, Interpretable Attention Heads In The Wild, December 2023
Gould, R., Ong, E., Ogden, G., and Conmy, A · 2023
Cited alongside, same era.
Finding Neurons in a Haystack: Case Studies with Sparse Probing, June 2023
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Cited alongside, same era.
Hanna, M., Liu, O., and Variengien, A · 2023
Cited alongside, same era.
dictionary_learning repository, 2023
Li, M., Marks, S., and Mueller, A · 2023
Cited alongside, same era.
Lieberum, T., Rahtz, M., Kramár, J., Nanda, N., Irving, G., Shah, R., and Mikulik, V · 2023
Cited alongside, same era.
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms, 2024
Hanna, M., Pezzelle, S., and Belinkov, Y · 2024
Closest in time.
He, Z., Ge, X., Tang, Q., Sun, T., Cheng, Q., and Qiu, X · 2024
Closest in time.
How to use and interpret activation patching
Heimersheim, S. and Nanda, N · 2024
Closest in time.
Attention SAEs Scale to GPT-2 Small, 2024
Kissane, C., Krzyzanowski, R., Conmy, A., and Nanda, N · 2024
Closest in time.
AtP*: An efficient and scalable method for localizing LLM behaviour to components, 2024
Kramár, J., Lieberum, T., Shah, R., and Nanda, N · 2024
Closest in time.
Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A · 2024
Closest in time.
Attribution Patching: Activation Patching At Industrial Scale, 2024
Neel Nanda · 2024
Closest in time.
Improving Dictionary Learning with Gated Sparse Autoencoders, April 2024
Rajamanoharan, S., Conmy, A., Smith, L., Lieberum, T., Varma, V., Kramár, J., Shah, R., and Nanda, N · 2024
Closest in time.
Predicting Future Activations, January 2024a
Templeton, A., Batson, J., Jermyn, A., and Olah, C · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Closest in time.
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods, January 2024
Zhang, F. and Nanda, N · 2024
Closest in time.