Fetching the paper…
Reading the bibliography…
Sparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling, 2023
Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., and van der Wal, O · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Earlier work this paper cites.
Sparse autoencoders find highly interpretable features in language models, 2023
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Earlier work this paper cites.
Finding neurons in a haystack: Case studies with sparse probing, 2023
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Earlier work this paper cites.
Open Source Replication & Commentary on Anthropic’s Dictionary Learning Paper, Oct 2023
Nanda, N · 2023
Earlier work this paper cites.
Circuits updates — august 2024
Anthropic Interpretability Team · 2024
Earlier work this paper cites.
Adaptive sparse allocation with mutual choice & feature choice sparse autoencoders, 2024
Ayonrinde, K · 2024
Earlier work this paper cites.
Evaluating open-source sparse autoencoders on disentangling factual knowledge in gpt-2 small, 2024
Chaudhary, M. and Geiger, A · 2024
Cited alongside, same era.
Not all language model features are linear, 2024
Engels, J., Michaud, E. J., Liao, I., Gurnee, W., and Tegmark, M · 2024
Cited alongside, same era.
Applying sparse autoencoders to unlearn knowledge in language models, 2024
Farrell, E., Lau, Y.-T., and Conmy, A · 2024
Cited alongside, same era.
Scaling and evaluating sparse autoencoders
Gao, L., Dupré la Tour, T., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J · 2024
Cited alongside, same era.
Gemma 2: Improving open language models at a practical size, 2024
Towards principled evaluations of sparse autoencoders for interpretability and control, 2024
Makelov, A., Lange, G., and Nanda, N · 2024
Later among the works it cites.
Efficient dictionary learning with switch sparse autoencoders
Mudide, A., Engels, J., Michaud, E. J., Tegmark, M., and Schroeder de Witt, C · 2024
Later among the works it cites.
Automatically interpreting millions of features in large language models, 2024
Paulo, G., Mallen, A., Juang, C., and Belrose, N · 2024
Later among the works it cites.
Prolu: A nonlinearity for sparse autoencoders, 2024
Taggart, G. M · 2024
Later among the works it cites.
Sage: Scalable ground truth evaluations for large sparse autoencoders, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gemma Team, Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., Ferret, J., Liu, P., Tafti, P., Friesen, A., Casbon, M., Ramos, S., Kumar, R., Lan, C. L., Jerome, S., Tsitsulin, A., Vieillard, N., Stanczyk, P., Girgin, S., Momchev, N., Hoffman, M., Thakoor, S., Grill, J.-B., Neyshabur, B., Bachem, O., Walton, A., Severyn, A., Parrish, A., Ahmad, A., Hutchison, A., Abdagic, A., Carl, A., Shen, A., Brock, A., Coenen, A., Laforge, A., Paterson, A., Bastian, B., Piot, B., Wu, B., Royal, B., Chen, C., Kumar, C., Perry, C., Welty, C., Choquette-Choo, C. A., Sinopalnikov, D., Weinberger, D., Vijaykumar, D., Rogozińska, D., Herbison, D., Bandy, E., Wang, E., Noland, E., Moreira, E., Senter, E., Eltyshev, E., Visin, F., Rasskin, G., Wei, G., Cameron, G., Martins, G., Hashemi, H., Klimczak-Plucińska, H., Batra, H., Dhand, H., Nardini, I., Mein, J., Zhou, J., Svensson, J., Stanway, J., Chan, J., Zhou, J. P., Carrasqueira, J., Iljazi, J., Becker, J., Fernandez, J., van Amersfoort, J., Gordon, J., Lipschultz, J., Newlan, J., yeong Ji, J., Mohamed, K., Badola, K., Black, K., Millican, K., McDonell, K., Nguyen, K., Sodhia, K., Greene, K., Sjoesund, L. L., Usui, L., Sifre, L., Heuermann, L., Lago, L., McNealus, L., Soares, L. B., Kilpatrick, L., Dixon, L., Martins, L., Reid, M., Singh, M., Iverson, M., Görner, M., Velloso, M., Wirth, M., Davidow, M., Miller, M., Rahtz, M., Watson, M., Risdal, M., Kazemi, M., Moynihan, M., Zhang, M., Kahng, M., Park, M., Rahman, M., Khatwani, M., Dao, N., Bardoliwalla, N., Devanathan, N., Dumai, N., Chauhan, N., Wahltinez, O., Botarda, P., Barnes, P., Barham, P., Michel, P., Jin, P., Georgiev, P., Culliton, P., Kuppala, P., Comanescu, R., Merhej, R., Jana, R., Rokni, R. A., Agarwal, R., Mullins, R., Saadat, S., Carthy, S. M., Cogan, S., Perrin, S., Arnold, S. M. R., Krause, S., Dai, S., Garg, S., Sheth, S., Ronstrom, S., Chan, S., Jordan, T., Yu, T., Eccles, T., Hennigan, T., Kocisky, T., Doshi, T., Jain, V., Yadav, V., Meshram, V., Dharmadhikari, V., Barkley, W., Wei, W., Ye, W., Han, W., Kwon, W., Xu, X., Shen, Z., Gong, Z., Wei, Z., Cotruta, V., Kirk, P., Rao, A., Giang, M., Peran, L., Warkentin, T., Collins, E., Barral, J., Ghahramani, Z., Hadsell, R., Sculley, D., Banks, J., Dragan, A., Petrov, S., Vinyals, O., Dean, J., Hassabis, D., Kavukcuoglu, K., Farabet, C., Buchatskaya, E., Borgeaud, S., Fiedel, N., Joulin, A., Kenealy, K., Dadashi, R., and Andreev, A · 2024
Cited alongside, same era.
Ravel: Evaluating interpretability methods on disentangling language model representations, 2024
Huang, J., Wu, Z., Potts, C., Geva, M., and Geiger, A · 2024
Cited alongside, same era.
Karvonen, A., Wright, B., Rager, C., Angell, R., Brinkmann, J., Smith, L., Verdun, C. M., Bau, D., and Marks, S · 2024
Cited alongside, same era.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024
Lieberum, T., Rajamanoharan, S., Conmy, A., Smith, L., Sonnerat, N., Varma, V., Kramár, J., Dragan, A., Shah, R., and Nanda, N · 2024
Cited alongside, same era.
Batchtopk: A simple improvement for topk-saes, 2024a
Bussmann, B., Leask, P., and Nanda, N
Cited in the paper.
Learning multi-level features with matryoshka saes, December 19 2024b
Bussmann, B., Leask, P., and Nanda, N
Cited in the paper.
Showing sae latents are not atomic using meta-saes, 2024c
Bussmann, B., Pearce, M., Leask, P., Bloom, J. I., Sharkey, L., and Nanda, N
Cited in the paper.
A is for absorption: Studying feature splitting and absorption in sparse autoencoders, 2024a
Chanin, D., Wilken-Smith, J., Dulka, T., Bhatnagar, H., and Bloom, J
Cited in the paper.
Venhoff, C., Calinescu, A., Torr, P., and de Witt, C. S · 2024
Later among the works it cites.
Training sparse autoencoders
Anthropic Interpretability Team · 2025
Closest in time.
Sparse autoencoders can interpret randomly initialized transformers, 2025
Heap, T., Lawson, T., Farnik, L., and Aitchison, L · 2025
Closest in time.