Fetching the paper…
Reading the bibliography…
Sparse autoencoders (SAEs) extract human-interpretable features from deep neural networks by transforming their activations into a sparse, higher dimensional latent space, and then reconstructing the activations from these latents.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A · 2018
Earlier work this paper cites.
From hard to soft: Understanding deep network nonlinearities via vector quantization and statistical inference
Balestriero, R. and Baraniuk, R · 2019
Earlier work this paper cites.
Openwebtext corpus
Gokaslan, A., Cohen, V., Pavlick, E., and Tellex, S · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
Interpreting neural networks through the polytope lens
Black, S., Sharkey, L., Grinsztajn, L., Winsor, E., Braun, D., Merizian, J., Parker, K., Guevara, C. R., Millidge, B., Alfour, G., and Leahy, C · 2022
Earlier work this paper cites.
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., et al · 2022
Earlier work this paper cites.
The singular value decompositions of transformer weight matrices are highly interpretable
Millidge, B. and Black, S · 2022
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Earlier work this paper cites.
Language models can explain neurons in language models
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N. L., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Earlier work this paper cites.
Redpajama: an open dataset for training large language models, 2023
Computer, T · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Cited alongside, same era.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Cited alongside, same era.
dictionary_learning repository, 2023
Li, M., Marks, S., and Mueller, A · 2023
Cited alongside, same era.
Interpretability as compression: Reconsidering sae explanations of neural activations with mdl-saes
Ayonrinde, K., Pearce, M. T., and Sharkey, L · 2024
Cited alongside, same era.
Universal neurons in gpt2 language models
Gurnee, W., Horsley, T., Guo, Z. C., Kheirkhah, T. R., Sun, Q., Hathaway, W., Nanda, N., and Bertsimas, D · 2024
Later among the works it cites.
Ghost grads: An improvement on resampling, 2024
Jermyn, A. and Templeton, A · 2024
Later among the works it cites.
Open source automated interpretability for sparse autoencoder features, July 2024
Juang, C., Paulo, G., Drori, J., and Belrose, N · 2024
Later among the works it cites.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A · 2024
Later among the works it cites.
Matryoshka sparse autoencoders
Nabeshima, N · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mechanistic permutability: Match features across layers
Balagansky, N., Maksimov, I., and Gavrilov, D · 2024
Cited alongside, same era.
Learning multi-level features with matryoshka saes
Bussman, B., Leask, P., and Nanda, N · 2024
Cited alongside, same era.
Bussmann, B., Leask, P., and Nanda, N · 2024
Cited alongside, same era.
A is for absorption: Studying feature splitting and absorption in sparse autoencoders
Chanin, D., Wilken-Smith, J., Dulka, T., Bhatnagar, H., and Bloom, J · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Transcoders find interpretable llm feature circuits
Dunefsky, J., Chlenski, P., and Nanda, N · 2024
Cited alongside, same era.
Decomposing the dark matter of sparse autoencoders
Engels, J., Riggs, L., and Tegmark, M · 2024
Cited alongside, same era.
Paulo, G., Mallen, A., Juang, C., and Belrose, N · 2024
Later among the works it cites.
The fineweb datasets: Decanting the web for the finest text data at scale
Penedo, G., Kydlíček, H., allal, L. B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L. V., and Wolf, T · 2024
Later among the works it cites.
Jumping ahead: Improving reconstruction fidelity with jumprelu sparse autoencoders
Rajamanoharan, S., Lieberum, T., Sonnerat, N., Conmy, A., Varma, V., Kramár, J., and Nanda, N · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., et al · 2024
Later among the works it cites.
Predicting future activations
Templeton, A., Batson, J., Jermyn, A., and Olah, C · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Saebench: A comprehensive benchmark for sparse autoencoders, 2024
Karvonen, A., Rager, C., Lin, J., Tigges, C., Bloom, J., Chanin, D., Lau, Y.-T., Farrell, E., Conmy, A., McDougall, C., Ayonrinde, K., Wearden, M., Marks, S., and Nanda, N · 2025
Closest in time.