Fetching the paper…
Reading the bibliography…
Sparse autoencoders (SAEs) are a popular method for interpreting concepts represented in large language model (LLM) activations.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Pang, B. and Lee, L · 2005
Earlier work this paper cites.
Revealing representational content with pattern-information fmri—an introductory guide
Mur, M., Bandettini, P. A., and Kriegeskorte, N · 2009
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G · 2016
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Chen, T. and Guestrin, C · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Davidson, T., Warmsley, D., Macy, M., and Weber, I · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F. Liu, M. G · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A · 2018
Earlier work this paper cites.
“going on a vacation” takes longer than “going for a walk”: A study of temporal commonsense understanding
Ben Zhou, Daniel Khashabi, Q. N. and Roth, D · 2019
Earlier work this paper cites.
”quartz: An open-domain dataset of qualitative relationship questions”
Tafjord, O., Gardner, M., Lin, K., and Clark, P · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Earlier work this paper cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Getting closer to AI complete question answering: A set of prerequisite real tasks
Rogers, A., Kovaleva, O., Downey, M., and Rumshisky, A · 2020
Earlier work this paper cites.
An interpretability illusion for bert
Bolukbasi, T., Pearce, A., Yuan, A., Coenen, A., Reif, E., Viégas, F., and Wattenberg, M · 2021
Earlier work this paper cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Elazar, Y., Ravfogel, S., Jacovi, A., and Goldberg, Y · 2021
Earlier work this paper cites.
Aligning ai with shared human values
Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Yun, Z., Chen, Y., Olshausen, B. A., and LeCun, Y · 2021
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Earlier work this paper cites.
Commonsenseqa 2.0: Exposing the limits of ai through gamification
Talmor, A., Yoran, O., Bras, R. L., Bhagavatula, C., Goldberg, Y., Choi, Y., and Berant, J · 2022
Earlier work this paper cites.
Language models can explain neurons in language models
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Earlier work this paper cites.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Earlier work this paper cites.
Machine-generated text detection using deep learning, 2023
Gaggar, R., Bhagchandani, A., and Oza, H · 2023
Earlier work this paper cites.
Language models represent space and time
Gurnee, W. and Tegmark, M · 2023
Earlier work this paper cites.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Cited alongside, same era.
Neuronpedia: Interactive reference and tooling for analyzing neural networks, 2023
Lin, J · 2023
Cited alongside, same era.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation, 2023
Lin, Z., Wang, Z., Tong, Y., Wang, Y., Guo, Y., Wang, Y., and Shang, J · 2023
Cited alongside, same era.
Marks, S. and Tegmark, M · 2023
Cited alongside, same era.
Emergent linear representations in world models of self-supervised sequence models
Nanda, N., Lee, A., and Wattenberg, M · 2023
Cited alongside, same era.
Gemma 2: Improving open language models at a practical size, 2024
Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., et al · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Later among the works it cites.
Spam Text Message Classification — kaggle.com
AI, T. and Ishii, D · 2025
Closest in time.
allenai/basic_arithmetic · Datasets at Hugging Face — huggingface.co
AllenAI · 2025
Closest in time.
Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku, 2024
Anthropic · 2025
Closest in time.
Using dictionary learning features as classifiers, October 2024a
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Representation engineering: A top-down approach to ai transparency
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al · 2023
Cited alongside, same era.
Mechanistic interpretability for ai safety–a review
Bereska, L. and Gavves, E · 2024
Cited alongside, same era.
Identifying functionally important features with end-to-end sparse dictionary learning, 2024
Braun, D., Taylor, J., Goldowsky-Dill, N., and Sharkey, L · 2024
Cited alongside, same era.
Using dictionary learning features as classifiers, October 2024b
Bricken, T., Marcus, J., Mishra-Sharma, S., Tong, M., Perez, E., Sharma, M., Rivoire, K., and Henighan, T · 2024
Cited alongside, same era.
Improving steering vectors by targeting sparse autoencoder features
Chalnev, S., Siu, M., and Conmy, A · 2024
Cited alongside, same era.
A is for absorption: Studying feature splitting and absorption in sparse autoencoders
Chanin, D., Wilken-Smith, J., Dulka, T., Bhatnagar, H., and Bloom, J · 2024
Cited alongside, same era.
Evaluating open-source sparse autoencoders on disentangling factual knowledge in gpt-2 small
Chaudhary, M. and Geiger, A · 2024
Cited alongside, same era.
Bricken, T., Marcus, J., Mishra-Sharma, S., Tong, M., Perez, E., Sharma, M., Rivoire, K., and Henighan, T · 2025
Closest in time.
Learning multi-level features with matryoshka saes
Bussmann, B., Leask, P., and Nanda, N · 2025
Closest in time.
fake-and-real-news-dataset — kaggle.com
clmentbisaillon · 2025
Closest in time.
Circuits Updates - April 2024 — transformer-circuits.pub
Conerly, T., Templeton, A., Bricken, T., Marcus, J., and Henighan, T · 2025
Closest in time.
Text classification documentation — kaggle.com
Dublish, T · 2025
Closest in time.
Medical Text Dataset -Cancer Doc Classification — kaggle.com
Falgunipatel19 · 2025
Closest in time.
Sparse autoencoder features for classifications and transferability, 2025
Gallifant, J., Chen, S., Sasse, K., Aerts, H., Hartvigsen, T., and Bitterman, D. S · 2025
Closest in time.
AI Vs Human Text — kaggle.com
Gerami, S · 2025
Closest in time.
IT Service Ticket Classification Dataset — kaggle.com
Goh, A · 2025
Closest in time.
Emotion Detection from Text — kaggle.com
Gupta, P · 2025
Closest in time.
Sparse autoencoders can interpret randomly initialized transformers, 2025
Heap, T., Lawson, T., Farnik, L., and Aitchison, L · 2025
Closest in time.
Saebench: A comprehensive benchmark for sparse autoencoders, December 2024b
Karvonen, A., Rager, C., Lin, J., Tigges, C., Bloom, J., Chanin, D., Lau, Y.-T., Farrell, E., Conmy, A., McDougall, C., Ayonrinde, K., Wearden, M., Marks, S., and Nanda, N · 2025
Closest in time.
Medical Text — kaggle.com
Kasaraneni, C. K · 2025
Closest in time.
rkotari/clickbait · Datasets at Hugging Face — huggingface.co
Kotari, R · 2025
Closest in time.
Hello GPT-4o
OpenAI · 2025
Closest in time.
Introducing openai o1
OpenAI · 2025
Closest in time.
Open problems in mechanistic interpretability, 2025
Sharkey, L., Chughtai, B., Batson, J., Lindsey, J., Wu, J., Bushnaq, L., Goldowsky-Dill, N., Heimersheim, S., Ortega, A., Bloom, J., Biderman, S., Garriga-Alonso, A., Conmy, A., Nanda, N., Rumbelow, J., Wattenberg, M., Schoots, N., Miller, J., Michaud, E. J., Casper, S., Tegmark, M., Saunders, W., Bau, D., Todd, E., Geiger, A., Geva, M., Hoogland, J., Murfet, D., and McGrath, T · 2025
Closest in time.
Interpreting preference models w/ sparse autoencoders, July 2024
Smith, L. R. and Brinkmann, J · 2025
Closest in time.
Stathead: Your all-access ticket to the Sports Reference database. — Stathead.com — stathead.com
Stathead · 2025
Closest in time.
Axbench: Steering llms? even simple baselines outperform sparse autoencoders, 2025
Wu, Z., Arora, A., Geiger, A., Wang, Z., Huang, J., Jurafsky, D., Manning, C. D., and Potts, C · 2025
Closest in time.