Fetching the paper…
Reading the bibliography…
Sparse autoencoders (SAEs) are a useful tool for uncovering human-interpretable features in the activations of large language models (LLMs).
Linear algebraic structure of word senses, with applications to polysemy
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A · 2018
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., et al · 2022
Earlier work this paper cites.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S., Hayase, J., and Srinivasa, S · 2023
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N. L., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Earlier work this paper cites.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Earlier work this paper cites.
Interpretability dreams
Olah, C · 2023
Earlier work this paper cites.
Sparse autoencoders do not find canonical units of analysis
Anonymous · 2024
Cited alongside, same era.
Interpretability as compression: Reconsidering sae explanations of neural activations with mdl-saes
Ayonrinde, K., Pearce, M. T., and Sharkey, L · 2024
Cited alongside, same era.
Mechanistic permutability: Match features across layers
Balagansky, N., Maksimov, I., and Gavrilov, D · 2024
Cited alongside, same era.
Sae repository
Belrose, N · 2024
Cited alongside, same era.
Identifying functionally important features with end-to-end sparse dictionary learning, 2024
Braun, D., Taylor, J., Goldowsky-Dill, N., and Sharkey, L · 2024
Cited alongside, same era.
Scaling and evaluating sparse autoencoders
Gao, L., la Tour, T. D., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J · 2024
Later among the works it cites.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024
Lieberum, T., Rajamanoharan, S., Conmy, A., Smith, L., Sonnerat, N., Varma, V., Kramár, J., Dragan, A., Shah, R., and Nanda, N · 2024
Later among the works it cites.
Enhancing neural network interpretability with feature-aligned sparse autoencoders
Marks, L., Paren, A., Krueger, D., and Barez, F · 2024
Later among the works it cites.
Automatically interpreting millions of features in large language models
Paulo, G., Mallen, A., Juang, C., and Belrose, N · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chanin, D., Wilken-Smith, J., Dulka, T., Bhatnagar, H., and Bloom, J · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Not all language model features are linear
Engels, J., Michaud, E. J., Liao, I., Gurnee, W., and Tegmark, M · 2024
Cited alongside, same era.
Rajamanoharan, S., Conmy, A., Smith, L., Lieberum, T., Varma, V., Kramár, J., Shah, R., and Nanda, N · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Later among the works it cites.
Saebench: A comprehensive benchmark for sparse autoencoders, 2024
Karvonen, A., Rager, C., Lin, J., Tigges, C., Bloom, J., Chanin, D., Lau, Y.-T., Farrell, E., Conmy, A., McDougall, C., Ayonrinde, K., Wearden, M., Marks, S., and Nanda, N · 2025
Closest in time.
The strong feature hypothesis could be wrong, 2024
Smith, L · 2025
Closest in time.