Fetching the paper…
Reading the bibliography…
In recent years many methods have been developed to understand the internal workings of neural networks, often by describing the function of individual neurons in the model.
Wordnet: a lexical database for english
Miller, G. A · 1995
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Erhan, D., Bengio, Y., Courville, A., and Vincent, P · 2009
Earlier work this paper cites.
Object detectors emerge in deep scene cnns
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
”why should i trust you?” explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes, 2018
Alain, G. and Bengio, Y · 2018
Earlier work this paper cites.
Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks
Fong, R. and Vedaldi, A · 2018
Earlier work this paper cites.
Understanding the role of individual units in a deep neural network
Bau, D., Zhu, J.-Y., Strobelt, H., Lapedriza, A., Zhou, B., and Torralba, A · 2020
Earlier work this paper cites.
Compositional explanations of neurons
Mu, J. and Andreas, J · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Multimodal neurons in artificial neural networks
Goh, G., Cammarata, N., Voss, C., Carter, S., Petrov, M., Schubert, L., Radford, A., and Olah, C · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Leveraging sparse linear layers for debuggable deep networks
Wong, E., Santurkar, S., and Madry, A · 2021
Cited alongside, same era.
Natural language descriptions of deep visual features
Hernandez, E., Schwettmann, S., Bau, D., Bagashvili, T., Torralba, A., and Andreas, J · 2022
Cited alongside, same era.
Finding neurons in a haystack: Case studies with sparse probing, 2023
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Later among the works it cites.
Identifying interpretable subspaces in image representations
Kalibhat, N., Bhardwaj, S., Bruss, C. B., Firooz, H., Sanjabi, M., and Feizi, S · 2023
Later among the works it cites.
The importance of prompt tuning for automated neuron explanations
Lee, J., Oikarinen, T., Chatha, A., Chang, K.-C., Chen, Y., and Weng, T.-W · 2023
Later among the works it cites.
Adversarial attacks on the interpretation of neuron activation maximization, 2023
Nanfack, G., Fulleringer, A., Marty, J., Eickenberg, M., and Belilovsky, E · 2023
Later among the works it cites.
Clip-dissect: Automatic description of neuron representations in deep vision networks
Oikarinen, T. and Weng, T.-W · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neuron-level interpretation of deep nlp models: A survey
Sajjad, H., Durrani, N., and Dalvi, F · 2022
Cited alongside, same era.
Language models can explain neurons in language models
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Cited alongside, same era.
Labeling neural representations with inverse recognition
Bykov, K., Kopf, L., Nakajima, S., Kloft, M., and Höhne, M. M · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models, 2023
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Cited alongside, same era.
Privileged bases in the transformer residual stream, 2023
Elhage, N., Lasenby, R., and Olah, C · 2023
Cited alongside, same era.
Don’t trust your eyes: on the (un)reliability of feature visualizations, 2023
Geirhos, R., Zimmermann, R. S., Bilodeau, B., Brendel, W., and Kim, B · 2023
Cited alongside, same era.
Rosa, B. L., Gilpin, L. H., and Capobianco, R · 2023
Later among the works it cites.
Corrupting neuron explanations of deep visual features
Srivastava, D., Oikarinen, T., and Weng, T.-W · 2023
Later among the works it cites.
Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip
Yu, Q., He, J., Deng, X., Shen, X., and Chen, L.-C · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L · 2023
Later among the works it cites.
Scale alone does not improve mechanistic interpretability in vision models
Zimmermann, R. S., Klein, T., and Brendel, W · 2023
Later among the works it cites.
Describe-and-dissect: Interpreting neurons in vision networks with language models, 2024
Bai, N., Iyer, R. A., Oikarinen, T., and Weng, T.-W · 2024
Closest in time.