Fetching the paper…
Reading the bibliography…
While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Reducing the Dimensionality of Data with Neural Networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
A bayesian hierarchical topic model for political texts: Measuring expressed agendas in senate press releases
J. Grimmer · 2010
Earlier work this paper cites.
To Explain or to Predict?
G. Shmueli · 2010
Earlier work this paper cites.
The importance of encoding versus training with sparse coding and vector quantization
A. Coates and A. Y. Ng · 2011
Earlier work this paper cites.
K-Sparse Autoencoders, Mar. 2014
A. Makhzani and B. Frey · 2014
Earlier work this paper cites.
Network Dissection: Quantifying Interpretability of Deep Visual Representations, Apr. 2017
D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba · 2017
Earlier work this paper cites.
Probing Classifiers: Promises, Shortcomings, and Advances
Y. Belinkov · 2017
Earlier work this paper cites.
Prediction and explanation in social systems
J. M. Hofman, A. Sharma, and D. J. Watts · 2017
Earlier work this paper cites.
Word embeddings quantify 100 years of gender and ethnic stereotypes
N. Garg, L. Schiebinger, D. Jurafsky, and J. Zou · 2018
Earlier work this paper cites.
Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study
J. R. Zech, M. A. Badgeley, M. Liu, A. B. Costa, J. J. Titano, and E. K. Oermann · 2018
Earlier work this paper cites.
Text as data
M. Gentzkow, B. Kelly, and M. Taddy · 2019
Earlier work this paper cites.
Clinicalbert: Modeling clinical notes and predicting hospital readmission
K. Huang, J. Altosaar, and R. Ranganath · 2019
Earlier work this paper cites.
Faithful and customizable explanations of black box models
H. Lakkaraju, E. Kamar, R. Caruana, and J. Leskovec · 2019
Earlier work this paper cites.
Linguistic Knowledge and Transferability of Contextual Representations, Apr. 2019
N. F. Liu, M. Gardner, Y. Belinkov, M. E. Peters, and N. A. Smith · 2019
Earlier work this paper cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
C. Rudin · 2019
Earlier work this paper cites.
Concept bottleneck models
P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, and P. Liang · 2020
Earlier work this paper cites.
Data leakage in health outcomes prediction with machine learning. Comment on “Prediction of incident hypertension within the next year: prospective study using statewide electronic health records and machine learning”
A. Chiavegatto Filho, A. F. D. M. Batista, and H. G. Dos Santos · 2021
Earlier work this paper cites.
Integrating explanation and prediction in computational social science
J. M. Hofman, D. J. Watts, S. Athey, F. Garip, T. L. Griffiths, J. Kleinberg, H. Margetts, S. Mullainathan, M. J. Salganik, S. Vazire, et al · 2021
Earlier work this paper cites.
Gender and representation bias in GPT-3 generated stories
L. Lucy and D. Bamman · 2021
Earlier work this paper cites.
Epic’s sepsis algorithm is going off the rails in the real world. the use of these variables may explain why
C. Ross · 2021
Earlier work this paper cites.
Computational analysis of 140 years of US political speeches reveals more positive but increasingly polarized framing of immigration
D. Card, S. Chang, C. Becker, J. Mendelsohn, R. Voigt, L. Boustan, R. Abramitzky, and D. Jurafsky · 2022
Earlier work this paper cites.
Toy Models of Superposition, Sept. 2022
N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, R. Grosse, S. McCandlish, J. Kaplan, D. Amodei, M. Wattenberg, and C. Olah · 2022
Cited alongside, same era.
Measuring the completeness of economic models
D. Fudenberg, J. Kleinberg, A. Liang, and S. Mullainathan · 2022
Cited alongside, same era.
Ai recognition of patient race in medical imaging: a modelling study
J. W. Gichoya, I. Banerjee, A. R. Bhimireddy, J. L. Burns, L. A. Celi, L.-C. Chen, R. Correa, N. Dullerud, M. Ghassemi, S.-C. Huang, et al · 2022
Cited alongside, same era.
Natural Language Descriptions of Deep Visual Features, Apr. 2022
E. Hernandez, S. Schwettmann, D. Bau, T. Bagashvili, A. Torralba, and J. Andreas · 2022
Cited alongside, same era.
Language models can explain neurons in language models
S. Bills, N. Cammarata, D. Mossing, H. Tillman, L. Gao, G. Goh, I. Sutskever, J. Leike, J. Wu, and W. Saunders · 2023
Cited alongside, same era.
Disentangling Dense Embeddings with Sparse Autoencoders, Aug. 2024
C. O’Neill, C. Ye, K. Iyer, and J. F. Wu · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet. transformer circuits thread, 2024
A. Templeton, T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones, et al · 2024
Later among the works it cites.
Tipping the balance: Predictive algorithms and institutional decision-making in context
S. Zhang · 2024
Later among the works it cites.
From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language Models, Feb. 2025
E. Adams, L. Bai, M. Lee, Y. Yu, and M. AlQuraishi · 2025
Closest in time.
SAEs are good for steering–if you select the right features
D. Arad, A. Mueller, and Y. Belinkov · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, et al · 2023
Cited alongside, same era.
Sparse Autoencoders Find Highly Interpretable Features in Language Models, Oct. 2023
H. Cunningham, A. Ewart, L. Riggs, R. Huben, and L. Sharkey · 2023
Cited alongside, same era.
Characterization of Stigmatizing Language in Medical Records
K. Harrigian, A. Zirikly, B. Chee, A. Ahmad, A. Links, S. Saha, M. C. Beach, and M. Dredze · 2023
Cited alongside, same era.
Leakage and the reproducibility crisis in machine-learning-based science
S. Kapoor and A. Narayanan · 2023
Cited alongside, same era.
Negativity drives online news consumption
C. E. Robertson, N. Pröllochs, K. Schwarzenegger, P. Pärnamets, J. J. Van Bavel, and S. Feuerriegel · 2023
Cited alongside, same era.
Golden Gate Claude, 2024
Anthropic · 2024
Cited alongside, same era.
Scaling automatic neuron description, October 2024
D. Choi, V. Huang, K. Meng, D. D. Johnson, J. Steinhardt, and S. Schwettmann · 2024
Cited alongside, same era.
Genome modeling and design across all domains of life with Evo 2, Feb. 2025
G. Brixi, M. G. Durrant, J. Ku, M. Poli, G. Brockman, D. Chang, G. A. Gonzalez, S. H. King, D. B. Li, A. T. Merchant, M. Naghipourfar, E. Nguyen, C. Ricci-Tam, D. W. Romero, G. Sun, A. Taghibakshi, A. Vorontsov, B. Yang, M. Deng, L. Gorton, N. Nguyen, N. K. Wang, E. Adams, S. A. Baccus, S. Dillmann, S. Ermon, D. Guo, R. Ilango, K. Janik, A. X. Lu, R. Mehta, M. R. K. Mofrad, M. Y. Ng, J. Pannu, C. Ré, J. C. Schmok, J. S. John, J. Sullivan, K. Zhu, G. Zynda, D. Balsam, P. Collison, A. B. Costa, T. Hernandez-Boussard, E. Ho, M.-Y. Liu, T. McGrath, K. Powell, D. P. Burke, H. Goodarzi, P. D. Hsu, and B. L. Hie · 2025
Closest in time.
Learning Multi-Level Features with Matryoshka Sparse Autoencoders, Mar. 2025
B. Bussmann, N. Nabeshima, A. Karvonen, and N. Nanda · 2025
Closest in time.
Saeuron: Interpretable concept unlearning in diffusion models with sparse autoencoders
B. Cywiński and K. Deja · 2025
Closest in time.
Aggregated individual reporting for post-deployment evaluation
J. Dai, I. D. Raji, B. Recht, and I. Y. Chen · 2025
Closest in time.
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit, Dec. 2025
N. Jiang, X. Sun, L. Dunlap, L. Smith, and N. Nanda · 2025
Closest in time.
Are sparse autoencoders useful? a case study in sparse probing
S. Kantamneni, J. Engels, S. Rajamanoharan, M. Tegmark, and N. Nanda · 2025
Closest in time.
On the biology of a large language model
J. Lindsey, W. Gurnee, E. Ameisen, B. Chen, A. Pearce, N. L. Turner, C. Citro, D. Abrahams, S. Carter, B. Hosmer, J. Marcus, M. Sklar, A. Templeton, T. Bricken, C. McDougall, H. Cunningham, T. Henighan, A. Jermyn, A. Jones, A. Persic, Z. Qi, T. B. Thompson, S. Zimmerman, K. Rivoire, T. Conerly, C. Olah, and J. Batson · 2025
Closest in time.
Science in the Age of Algorithms
S. Mullainathan and A. Rambachan · 2025
Closest in time.
Deploying interpretability to production with rakuten: Sae probes for pii detection
N. Nguyen, M. Deng, D. Gala, K. Naruse, F. G. Virgo, M. Byun, D. Hazra, L. Gorton, D. Balsam, T. McGrath, M. Takei, and Y. Kaji · 2025
Closest in time.
Transcoders Beat Sparse Autoencoders for Interpretability, Feb. 2025
G. Paulo, S. Shabalin, and N. Belrose · 2025
Closest in time.
Open problems in mechanistic interpretability
L. Sharkey, B. Chughtai, J. Batson, J. Lindsey, J. Wu, L. Bushnaq, N. Goldowsky-Dill, S. Heimersheim, A. Ortega, J. Bloom, et al · 2025
Closest in time.
InterPLM: Discovering interpretable features in protein language models via sparse autoencoders
E. Simon and J. Zou · 2025
Closest in time.
Negative results for sparse autoencoders on downstream tasks and deprioritising sae research (mechanistic interpretability team progress update), 2025
L. Smith, S. Rajamanoharan, A. Conmy, C. McDougall, J. Kramar, T. Lieberum, R. Shah, and N. Nanda · 2025
Closest in time.
Concept Bottleneck Large Language Models, Sept. 2025
C.-E. Sun, T. Oikarinen, B. Ustun, and T.-W. Weng · 2025
Closest in time.
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models, June 2025
L. Tjuatja and G. Neubig · 2025
Closest in time.
Axbench: Steering llms? even simple baselines outperform sparse autoencoders
Z. Wu, A. Arora, A. Geiger, Z. Wang, J. Huang, D. Jurafsky, C. D. Manning, and C. Potts · 2025
Closest in time.
A machine learning model using clinical notes to identify physician fatigue
C.-C. Hsu, Z. Obermeyer, and C. Tan · 2041
Closest in time.