Fetching the paper…
Reading the bibliography…
Prior interpretability research studying narrow distributions has preliminarily identified self-repair, a phenomena where if components in large language models are ablated, later components will change their behavior to compensate.
Highway and residual networks learn unrolled iterative estimation, 2017
Greff, K., Srivastava, R. K., and Schmidhuber, J · 2017
Earlier work this paper cites.
Residual connections encourage iterative inference, 2018
Jastrzębski, S., Arpit, D., Ballas, N., Verma, V., Che, T., and Bengio, Y · 2018
Earlier work this paper cites.
Universal transformers, 2019
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Łukasz Kaiser · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?, 2019
Michel, P., Levy, O., and Neubig, G · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Understanding the role of individual units in a deep neural network
Bau, D., Zhu, J.-Y., Strobelt, H., Lapedriza, A., Zhou, B., and Torralba, A · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Towards falsifiable interpretability research, 2020
Leavitt, M. L. and Morcos, A · 2020
Earlier work this paper cites.
interpreting gpt: the logit lens, 2020
nostalgebraist · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
Causal mediation analysis for interpreting neural nlp: The case of gender bias, 2020
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Sakenis, S., Huang, J., Singer, Y., and Shieber, S · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Causal scrubbing, a method for rigorously testing interpretability hypotheses
Chan, L., Garriga-Alonso, A., Goldwosky-Dill, N., Greenblatt, R., Nitishinskaya, J., Radhakrishnan, A., Shlegeris, B., and Thomas, N · 2022
Cited alongside, same era.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens, 2023
Belrose, N., Furman, Z., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J · 2023
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Finding neurons in a haystack: Case studies with sparse probing, 2023
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Later among the works it cites.
Residual stream norms grow exponentially over the forward pass, 2023
Heimersheim, S. and Turner, A · 2023
Later among the works it cites.
Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla, 2023
Lieberum, T., Rahtz, M., Kramár, J., Nanda, N., Irving, G., Shah, R., and Mikulik, V · 2023
Later among the works it cites.
Copy suppression: Comprehensively understanding an attention head, 2023
McDougall, C., Conmy, A., Rushing, C., McGrath, T., and Nanda, N · 2023
Later among the works it cites.
The hydra effect: Emergent self-repair in language model computations, 2023
McGrath, T., Rahtz, M., Kramar, J., Mikulik, V., and Legg, S · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., and Van Der Wal, O · 2023
Cited alongside, same era.
Language models can explain neurons in language models
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Cited alongside, same era.
Analyzing transformers in embedding space, 2023
Dar, G., Geva, M., Gupta, A., and Berant, J · 2023
Cited alongside, same era.
Localizing model behavior with path patching, 2023
Goldowsky-Dill, N., MacLeod, C., Sato, L., and Arora, A · 2023
Cited alongside, same era.
Successor heads: Recurring, interpretable attention heads in the wild, 2023
Gould, R., Ong, E., Ogden, G., and Conmy, A · 2023
Cited alongside, same era.
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
Neurons in large language models: Dead, n-gram, positional, 2023
Voita, E., Ferrando, J., and Nalmpantis, C · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2023
Later among the works it cites.
Universal neurons in gpt2 language models, 2024
Gurnee, W., Horsley, T., Guo, Z. C., Kheirkhah, T. R., Sun, Q., Hathaway, W., Nanda, N., and Bertsimas, D · 2024
Closest in time.