Rosetta neurons: Mining the common units in a model zoo
Dravid, A., Gandelsman, Y., Efros, A. A., and Shocher, A. (2023) · 1943
Earlier work this paper cites.
Analyzing redundancy in pretrained transformer models
Original
Dalvi, F., Sajjad, H., Durrani, N., and Belinkov, Y. (2020) · 2004
Earlier work this paper cites.
Levels of analysis for machine learning
Original
Hamrick, J. and Mohamed, S. (2020) · 2004
Earlier work this paper cites.
Vision: A computational investigation into the human representation and processing of visual information
Marr, D. (2010) · 2010
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Original
Geva, M., Schuster, R., Berant, J., and Levy, O. (2020) · 2012
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Original
Li, Y., Yosinski, J., Clune, J., Lipson, H., and Hopcroft, J. (2015) · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Original
Hendrycks, D. and Gimpel, K. (2016) · 2016
Earlier work this paper cites.
Convergent Learning: Do different neural networks learn the same representations?
Li, Y., Yosinski, J., Clune, J., Lipson, H., and Hopcroft, J. (2016) · 2016
Earlier work this paper cites.
Multifaceted feature visualization: Uncovering the different types of features learned by each neuron in deep neural networks
Nguyen, A., Yosinski, J., and Clune, J. (2016) · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Original
Doshi-Velez, F. and Kim, B. (2017) · 2017
Earlier work this paper cites.
Residual connections encourage iterative inference
Original
Jastrzębski, S., Arpit, D., Ballas, N., Verma, V., Che, T., and Bengio, Y. (2017) · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Original
Radford, A., Jozefowicz, R., and Sutskever, I. (2017) · 2017
Earlier work this paper cites.
SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J. (2017) · 2017
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
Original
Bau, A., Belinkov, Y., Sajjad, H., Durrani, N., Dalvi, F., and Glass, J. (2018) · 2018
Earlier work this paper cites.
Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change
Hamilton, W. L., Leskovec, J., and Jurafsky, D. (2018) · 2018
Earlier work this paper cites.
Geometry Score: A Method For Comparing Generative Adversarial Networks
Khrulkov, V. and Oseledets, I. (2018) · 2018
Earlier work this paper cites.
Insights on representational similarity in neural networks with canonical correlation
Morcos, A. S., Raghu, M., and Bengio, S. (2018) · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al. (2018) · 2018
Earlier work this paper cites.
Interactive supercomputing on 40,000 cores for machine learning and data analysis
Reuther, A., Kepner, J., Byun, C., Samsi, S., Arcand, W., Bestor, D., Bergeron, B., Gadepally, V., Houle, M., Hubbell, M., Jones, M., Klein, A., Milechin, L., Mullen, J., Prout, A., Rosa, A., Yee, C., and Michaleas, P. (2018) · 2018
Earlier work this paper cites.
Towards Understanding Learning Representations: To What Extent Do Different Neural Networks Learn the Same Representation
Wang, L., Hu, L., Gu, J., Wu, Y., Hu, Z., He, K., and Hopcroft, J. (2018) · 2018
Earlier work this paper cites.
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models
Dalvi, F., Durrani, N., Sajjad, H., Belinkov, Y., Bau, A., and Glass, J. (2019) · 2019
Earlier work this paper cites.
On interpretability and feature representations: an analysis of the sentiment neuron
Donnelly, J. and Roegiest, A. (2019) · 2019
Earlier work this paper cites.
Similarity of neural network representations revisited
Kornblith, S., Norouzi, M., Lee, H., and Hinton, G. (2019) · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. (2019) · 2019
Earlier work this paper cites.
What part of the neural network does this? understanding lstms by measuring and dissecting neurons
Xin, J., Lin, J., and Yu, Y. (2019) · 2019
Earlier work this paper cites.
Understanding the role of individual units in a deep neural network
Bau, D., Zhu, J.-Y., Strobelt, H., Lapedriza, A., Zhou, B., and Torralba, A. (2020) · 2020
Earlier work this paper cites.
Curve circuits
Cammarata, N., Goh, G., Carter, S., Voss, C., Schubert, L., and Olah, C. (2021) · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Original
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. (2020) · 2020
Earlier work this paper cites.
spacy: Industrial-strength natural language processing in python
Honnibal, M., Montani, I., Van Landeghem, S., and Boyd, A. (2020) · 2020
Earlier work this paper cites.
Inter-layer Information Similarity Assessment of Deep Neural Networks Via Topological Similarity and Persistence Analysis of Data Neighbour Dynamics
Hryniowski, A. and Wong, A. (2020) · 2020
Earlier work this paper cites.
Compositional explanations of neurons
Mu, J. and Andreas, J. (2020) · 2020
Earlier work this paper cites.
Interpreting gpt: The logit lens
Nostalgebraist (2020) · 2020
Earlier work this paper cites.
An overview of early vision in inceptionv1
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S. (2020a) · 2020
Earlier work this paper cites.
High-low frequency detectors
Schubert, L., Voss, C., Cammarata, N., Goh, G., and Olah, C. (2021a) · 2020
Earlier work this paper cites.
Similarity of Neural Networks with Gradients
Tang, S., Maddox, W. J., Dickens, C., Diethe, T., and Damianou, A. (2020) · 2020
Earlier work this paper cites.