Fetching the paper…
Reading the bibliography…
Superposition -- when a neural network represents more ``features'' than it has dimensions -- seems to pose a serious challenge to mechanistically interpreting current AI systems.
On a modification of chebyshev’s inequality and of the error formula of laplace
Bernstein, S · 1924
Earlier work this paper cites.
Principles of neurodynamics. perceptrons and the theory of brain mechanisms
Rosenblatt, F · 1961
Earlier work this paper cites.
Parallel distributed processing: explorations in the microstructure of cognition
Holyoak, K. J · 1987
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Fodor, J. A. and Pylyshyn, Z. W · 1988
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
A note on negligible functions
Bellare · 2002
Earlier work this paper cites.
Uniform approximation of functions with random bases
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories, September 2021
Geva, M., Schuster, R., Berant, J., and Levy, O · 2012
Earlier work this paper cites.
On the computational intractability of exact and approximate dictionary learning
Tillmann, A. M · 2014
Earlier work this paper cites.
Why neurons mix: high dimensionality for higher cognition
Fusi, S., Miller, E. K., and Rigotti, M · 2016
Earlier work this paper cites.
Multifaceted feature visualization: Uncovering the different types of features learned by each neuron in deep neural networks, 2016
Nguyen, A., Yosinski, J., and Clune, J · 2016
Earlier work this paper cites.
Feature visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Causal abstractions of neural networks
Geiger, A., Lu, H., Icard, T., and Potts, C · 2021
Cited alongside, same era.
Multimodal neurons in artificial neural networks
Goh, G., †, N. C., †, C. V., Carter, S., Petrov, M., Schubert, L., Radford, A., and Olah, C · 2021
Cited alongside, same era.
Spiking hyperdimensional network: Neuromorphic models integrated with memory-inspired framework
Locating and editing factual associations in gpt, 2023
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2023
Later among the works it cites.
Randomized numerical linear algebra: A perspective on the field with an eye to software
Murray, R., Demmel, J., Mahoney, M. W., Erichson, N. B., Melnichenko, M., Malik, O. A., Grigori, L., Luszczek, P., Dereziński, M., Lopes, M. E., et al · 2023
Later among the works it cites.
Toward transparent AI: A survey on interpreting the inner structures of deep neural networks, 2023
Räuker, T., Ho, A., Casper, S., and Hadfield-Menell, D · 2023
Later among the works it cites.
Codebook features: Sparse and discrete interpretability for neural networks
Tamkin, A., Taufeeque, M., and Goodman, N. D · 2023
Later among the works it cites.
Open source sparse autoencoders for all residual stream layers of GPT2 small
Bloom, J · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zou, Z., Alimohamadi, H., Imani, F., Kim, Y., and Imani, M · 2021
Cited alongside, same era.
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., et al · 2022
Cited alongside, same era.
Polysemanticity and capacity in neural networks
Scherlis, A., Sachan, K., Jermyn, A. S., Benton, J., and Shlegeris, B · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Cited alongside, same era.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Cited alongside, same era.
Identifying functionally important features with end-to-end sparse dictionary learning
Braun, D., Taylor, J., Goldowsky-Dill, N., and Sharkey, L · 2024
Closest in time.
Towards automated circuit discovery for mechanistic interpretability
Conmy, A., Mavor-Parker, A., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2024
Closest in time.
Interim research report: Activation plateaus and sensitive periods in transformer training, 2023
Heimersheim, S. and Mendel, J · 2024
Closest in time.
What’s up with llms representing xors of arbitrary features?, 2024
Marks, S · 2024
Closest in time.
Improving dictionary learning with gated sparse autoencoders
Rajamanoharan, S., Conmy, A., Smith, L., Lieberum, T., Varma, V., Kramár, J., Shah, R., and Nanda, N · 2024
Closest in time.
ProLU: A nonlinearity for sparse autoencoders
Taggart, G. M · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Closest in time.