Fetching the paper…
Reading the bibliography…
Mechanistic Interpretability (MI) promises a path toward fully understanding how neural networks make their predictions.
Über den zusammenhang des abschlusses der elektronengruppen im atom mit der komplexstruktur der spektren
Pauli, W · 1925
Earlier work this paper cites.
Zur theorie der kernmassen
Weizsäcker, C. F. v · 1935
Earlier work this paper cites.
Nuclear Physics A. Stationary States of Nuclei
Bethe, H. A. and Bacher, R. F · 1936
Earlier work this paper cites.
Saliency, scale and image description
Kadir, T. and Brady, M · 2001
Earlier work this paper cites.
Mutual influence of terms in a semi-empirical mass formula
Kirson, M. W · 2007
Earlier work this paper cites.
Interpreting principal component analyses of spatial population genetic variation
Novembre, J. and Stephens, M · 2008
Earlier work this paper cites.
Table of experimental nuclear ground state charge radii: An update
Angeli, I. and Marinova, K. P · 2011
Earlier work this paper cites.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
Feature visualization
Olah, C., Schubert, L., and Mordvintsev, A · 2017
Earlier work this paper cites.
Pca of high dimensional random walks with comparison to neural network training
Antognini, J. and Sohl-Dickstein, J · 2018
Earlier work this paper cites.
Understanding disentangling in β \beta -vae
Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., and Lerchner, A · 2018
Earlier work this paper cites.
Isolating sources of disentanglement in variational autoencoders
Chen, R. T., Li, X., Grosse, R. B., and Duvenaud, D. K · 2018
Earlier work this paper cites.
Towards a definition of disentangled representations
Higgins, I., Amos, D., Pfau, D., Racaniere, S., Matthey, L., Rezende, D., and Lerchner, A · 2018
Cited alongside, same era.
Disentangling by factorising
Kim, H. and Mnih, A · 2018
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes
Li, C., Farkhoor, H., Liu, R., and Yosinski, J · 2018
Cited alongside, same era.
Analysis of neuronal ensemble activity reveals the pitfalls and shortcomings of rotation dynamics
Lebedev, M. A., Ossadtchi, A., Mill, N. A., Urpí, N. A., Cervera, M. R., and Nicolelis, M. A · 2019
Cited alongside, same era.
Challenging common assumptions in the unsupervised learning of disentangled representations
Locatello, F., Bauer, S., Lucic, M., Raetsch, G., Gelly, S., Sch”olkopf, B., and Bachem, O · 2019
Cited alongside, same era.
WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models
Benchekroun, Y., Dervishi, M., Ibrahim, M., Gaya, J.-B., Martinet, X., Mialon, G., Scialom, T., Dupoux, E., Hupkes, D., and Vincent, P · 2023
Later among the works it cites.
Eight Things to Know about Large Language Models
Bowman, S. R · 2023
Later among the works it cites.
Interpretable machine learning for science with pysr and symbolicregression. jl
Cranmer, M · 2023
Later among the works it cites.
Discovery of a planar black hole mass scaling relation for spiral galaxies
Davis, B. L. and Jin, Z · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A., Gupta, S., and Zettlemoyer, L · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
The AME 2020 atomic mass evaluation (II). Tables, graphs and references
Wang, M., Huang, W. J., Kondev, F. G., Audi, G., and Naimi, S · 2021
Cited alongside, same era.
A survey on neural network interpretability
Zhang, Y., Tiňo, P., Leonardis, A., and Tang, K · 2021
Cited alongside, same era.
Hassid, M., Peng, H., Rotem, D., Kasai, J., Montero, I., Smith, N. A., and Schwartz, R · 2022
Cited alongside, same era.
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M · 2022
Cited alongside, same era.
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Later among the works it cites.
Language Models Represent Space and Time
Gurnee, W. and Tegmark, M · 2023
Later among the works it cites.
Rediscovering orbital mechanics with machine learning
Lemos, P., Jeffrey, N., Cranmer, M., Ho, S., and Battaglia, P · 2023
Later among the works it cites.
Interpretable machine learning methods applied to jet background subtraction in heavy ion collisions
Mengel, T., Steffanic, P., Hughes, C., da Silva, A. C. O., and Nattrass, C · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Later among the works it cites.
GPT4GEO: How a Language Model Sees the World’s Geography
Roberts, J., Lüddecke, T., Das, S., Han, K., and Albanie, S · 2023
Later among the works it cites.
Phantom oscillations in principal component analysis
Shinn, M · 2023
Later among the works it cites.
AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
Zhang, Q., Chen, M., Bukharin, A., Karampatziakis, N., He, P., Cheng, Y., Chen, W., and Zhao, T · 2023
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Zhong, Z., Liu, Z., Tegmark, M., and Andreas, J · 2023
Later among the works it cites.
Slicegpt: Compress large language models by deleting rows and columns, 2024
Ashkboos, S., Croci, M. L., do Nascimento, M. G., Hoefler, T., and Hensman, J · 2024
Closest in time.