Fetching the paper…
Reading the bibliography…
Prior work has shown the existence of contextual neurons in language models, including a neuron that activates on German text.
Direct and indirect effects
Pearl, J. (2001) · 2001
Earlier work this paper cites.
Metrics for evaluating 3d medical image segmentation: analysis, selection, and tool
Taha, A. A. and Hanbury, A. (2015) · 2015
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J. (2017) · 2017
Earlier work this paper cites.
A mathematical theory of semantic development in deep neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S. (2019) · 2019
Earlier work this paper cites.
Thread: Circuits
Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K. (2020) · 2020
Earlier work this paper cites.
Pretrained language model embryology: The birth of ALBERT
Chiang, C.-H., Huang, S.-F., and Lee, H.-y. (2020) · 2020
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Dodge, J., Ilharco, G., Schwartz, R., Farhadi, A., Hajishirzi, H., and Smith, N. (2020) · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C. (2020) · 2020
Earlier work this paper cites.
Towards falsifiable interpretability research
Leavitt, M. L. and Morcos, A. (2020) · 2020
Earlier work this paper cites.
Berts of a feather do not generalize together: Large variability in generalization across models with similar test set performance
McCoy, R. T., Min, J., and Linzen, T. (2020) · 2020
Earlier work this paper cites.
LSTMs compose—and Learn—Bottom-up
Saphra, N. and Lopez, A. (2020a) · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S. (2020) · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. (2021) · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O. (2021) · 2021
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Burns, C., Ye, H., Klein, D., and Steinhardt, J. (2022) · 2022
Cited alongside, same era.
Causal scrubbing, a method for rigorously testing interpretability hypotheses
Chan, L., Garriga-Alonso, A., Goldwosky-Dill, N., Greenblatt, R., Nitishinskaya, J., Radhakrishnan, A., Shlegeris, B., and Thomas, N. (2022) · 2022
Cited alongside, same era.
Softmax linear units
Elhage, N., Hume, T., Olsson, C., Nanda, N., Henighan, T., Johnston, S., ElShowk, S., Joseph, N., DasSarma, N., Mann, B., Hernandez, D., Askell, A., Ndousse, K., Jones, A., Drain, D., Chen, A., Bai, Y., Ganguli, D., Lovitt, L., Hatfield-Dodds, Z., Kernion, J., Conerly, T., Kravec, S., Fort, S., Kadavath, S., Jacobson, J., Tran-Johnson, E., Kaplan, J., Clark, J., Brown, T., McCandlish, S., Amodei, D., and Olah, C. (2022a) · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Geva, M., Caciularu, A., Wang, K. R., and Goldberg, Y. (2022) · 2022
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C. (2023) · 2023
Closest in time.
The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation
Chicco1, D. and Jurman, G. (2023) · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D. (2023) · 2023
Closest in time.
Linear connectivity reveals generalization strategies
Juneja, J., Bansal, R., Cho, K., Sedoc, J., and Saphra, N. (2023) · 2023
Closest in time.
The hydra effect: Emergent self-repair in language model computations
McGrath, T., Rahtz, M., Kramar, J., Mikulik, V., and Legg, S. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Transformerlens
Nanda, N. and Bloom, J. (2022) · 2022
Cited alongside, same era.
Mechanistic interpretability, variables, and the importance of interpretable bases. transformer circuits thread (june 27)
Olah, C. (2022) · 2022
Cited alongside, same era.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. (2022) · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J. (2022) · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. (2022) · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., and van der Wal, O. (2023) · 2023
Cited alongside, same era.
Language models can explain neurons in language models
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W. (2023) · 2023
Cited alongside, same era.
Closest in time.
Locating and editing factual associations in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y. (2023) · 2023
Closest in time.
The quantization model of neural scaling
Michaud, E. J., Liu, Z., Girit, U., and Tegmark, M. (2023) · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J. (2023) · 2023
Closest in time.
interpreting gpt: the logit lens
nostalgebraist (2020) · 2023
Closest in time.
On the special role of class-selective neurons in early training
Ranadive, O., Thakurdesai, N., Morcos, A. S., and Matthew Leavitt, S. D. (2023) · 2023
Closest in time.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Räuker, T., Ho, A., Casper, S., and Hadfield-Menell, D. (2023) · 2023
Closest in time.
Explaining grokking through circuit efficiency
Varma, V., Shah, R., Kenton, Z., Kramár, J., and Kumar, R. (2023) · 2023
Closest in time.