Fetching the paper…
Reading the bibliography…
A key goal of current mechanistic interpretability research in NLP is to find linear features (also called "feature vectors") for transformers: directions in activation space corresponding to concepts that are used by a given model in its computation.
The matrix cookbook, version 2012/11/15
Petersen, K. and Pedersen, M · 2012
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V., and Kalai, A. T · 2016
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., and sayres, R · 2018
Earlier work this paper cites.
The building blocks of interpretability
Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., and Mordvintsev, A · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
AllenNLP interpret: A framework for explaining predictions of NLP models
Wallace, E., Tuyls, J., Wang, J., Subramanian, S., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2020
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
Gender by Name
UCI Machine Learning Repository · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S · 2020
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2020
Earlier work this paper cites.
GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow, 2021
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S · 2021
Earlier work this paper cites.
mdmm, 2021
Crowson, K · 2021
Earlier work this paper cites.
Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals
Elazar, Y., Ravfogel, S., Jacovi, A., and Goldberg, Y · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Contrastive explanations for model interpretability
Jacovi, A., Swayamdipta, S., Ravfogel, S., Elazar, Y., Choi, Y., and Goldberg, Y · 2021
Cited alongside, same era.
Softmax linear units
Elhage, N., Hume, T., Olsson, C., Nanda, N., Henighan, T., Johnston, S., ElShowk, S., Joseph, N., DasSarma, N., Mann, B., Hernandez, D., Askell, A., Ndousse, K., Jones, A., Drain, D., Chen, A., Bai, Y., Ganguli, D., Lovitt, L., Hatfield-Dodds, Z., Kernion, J., Conerly, T., Kravec, S., Fort, S., Kadavath, S., Jacobson, J., Tran-Johnson, E., Kaplan, J., Clark, J., Brown, T., McCandlish, S., Amodei, D., and Olah, C · 2022
Cited alongside, same era.
A comprehensive mechanistic interpretability explainer & glossary, 2022
Nanda, N · 2022
Analyzing transformers in embedding space
Dar, G., Geva, M., Gupta, A., and Berant, J · 2023
Closest in time.
Localizing model behavior with path patching
Goldowsky-Dill, N., MacLeod, C., Sato, L., and Arora, A · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing, 2023
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Closest in time.
Linearity of relation decoding in transformer language models, 2023
Hernandez, E., Sharma, A. S., Haklay, T., Meng, K., Wattenberg, M., Andreas, J., Belinkov, Y., and Bau, D · 2023
Closest in time.
Uncovering intermediate variables in transformers using circuit probing, 2023
Lepori, M. A., Serre, T., and Pavlick, E · 2023
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mechanistic interpretability, variables, and the importance of interpretable bases, 2022
Olah, C · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Cited alongside, same era.
Re-examining layernorm
Winsor, E · 2022
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Cited alongside, same era.
On the expressivity role of layernorm in transformers’ attention
Brody, S., Alon, U., and Yahav, E · 2023
Cited alongside, same era.
Li, K., Patel, O., Viégas, F., Pfister, H., and Wattenberg, M · 2023
Closest in time.
Identifying a preliminary circuit for predicting gendered pronouns in gpt-2 small
Mathwin, C., Corlouer, G., Kran, E., Barez, F., and Nanda, N · 2023
Closest in time.
The singular value decompositions of transformer weight matrices are highly interpretable
Millidge, B. and Black, S · 2023
Closest in time.
Attribution patching: Activation patching at industrial scale
Nanda, N., Olah, C., Olsson, C., Elhage, N., and Tristan, H · 2023
Closest in time.
The linear representation hypothesis and the geometry of large language models, 2023
Park, K., Choe, Y. J., and Veitch, V · 2023
Closest in time.
Attribution patching outperforms automated circuit discovery, 2023
Syed, A., Rager, C., and Conmy, A · 2023
Closest in time.
Linear representations of sentiment in large language models
Tigges, C., Hollinsworth, O. J., Geiger, A., and Nanda, N · 2023
Closest in time.
Characterizing mechanisms for factual recall in language models, 2023
Yu, Q., Merullo, J., and Pavlick, E · 2023
Closest in time.
Interpretability at scale: Identifying causal mechanisms in alpaca, 2024
Wu, Z., Geiger, A., Potts, C., and Goodman, N. D · 2024
Closest in time.