Fetching the paper…
Reading the bibliography…
Sparse autoencoders (SAEs) have been successfully used to discover sparse and human-interpretable representations of the latent activations of LLMs.
Jumprelu: A retrofit defense strategy for adversarial attacks, 2019
Erichson, N. B., Yao, Z., and Mahoney, M. W · 1904
Earlier work this paper cites.
Glu variants improve transformer, 2020
Shazeer, N · 2002
Earlier work this paper cites.
Language modeling with gated convolutional networks, 2017
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization, January 2017
Kingma, D. P. and Ba, J · 2017
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Thread: Circuits
Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K · 2020
Earlier work this paper cites.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling, December 2020
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Zoom In: An Introduction to Circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors
Yun, Z., Chen, Y., Olshausen, B., and LeCun, Y · 2021
Earlier work this paper cites.
Locating and Editing Factual Associations in GPT
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Earlier work this paper cites.
In-context Learning and Induction Heads, September 2022
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Earlier work this paper cites.
Taking features out of superposition with sparse autoencoders, December 2022
Sharkey, L., Braun, D., and Millidge, B · 2022
Earlier work this paper cites.
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small, November 2022
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Earlier work this paper cites.
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., and Wal, O. V. D · 2023
Earlier work this paper cites.
Language models can explain neurons in language models, May 2023
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Earlier work this paper cites.
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning, 2023
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., and Askell, A · 2023
Cited alongside, same era.
Towards Automated Circuit Discovery for Mechanistic Interpretability
Conmy, A., Mavor-Parker, A., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Cited alongside, same era.
Sparse Autoencoders Find Highly Interpretable Features in Language Models, October 2023
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Cited alongside, same era.
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Hanna, M., Liu, O., and Variengien, A · 2023
Cited alongside, same era.
Attribution Patching: Activation Patching At Industrial Scale, February 2023
Nanda, N · 2023
Cited alongside, same era.
Sparse autoencoders reveal universal feature spaces across large language models, 2024
Lan, M., Torr, P., Meek, A., Khakzar, A., Krueger, D., and Barez, F · 2024
Later among the works it cites.
Residual Stream Analysis with Multi-Layer SAEs, October 2024
Lawson, T., Farnik, L., Houghton, C., and Aitchison, L · 2024
Later among the works it cites.
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2, August 2024
Lieberum, T., Rajamanoharan, S., Conmy, A., Smith, L., Sonnerat, N., Varma, V., Kramár, J., Dragan, A., Shah, R., and Nanda, N · 2024
Later among the works it cites.
Sparse Autoencoders Match Supervised Features for Model Steering on the IOI Task
Makelov, A · 2024
Later among the works it cites.
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models, March 2024
Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
Zhang, F. and Nanda, N · 2023
Cited alongside, same era.
Decomposing and interpreting image representations via text in vits beyond clip, 2024
Balasubramanian, S., Basu, S., and Feizi, S · 2024
Cited alongside, same era.
Evolution of sae features across layers in llms, 2024
Balcells, D., Lerner, B., Oesterle, M., Ucar, E., and Heimersheim, S · 2024
Cited alongside, same era.
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning, May 2024
Braun, D., Taylor, J., Goldowsky-Dill, N., and Sharkey, L · 2024
Cited alongside, same era.
Scaling Automatic Neuron Description, October 2024
Choi, D., Huang, V., Meng, K., Johnson, D. D., Steinhardt, J., and Schwettmann, S · 2024
Cited alongside, same era.
Transcoders Find Interpretable LLM Feature Circuits, June 2024
Dunefsky, J., Chlenski, P., and Nanda, N · 2024
Cited alongside, same era.
Applying sparse autoencoders to unlearn knowledge in language models, 2024
Farrell, E., Lau, Y.-T., and Conmy, A · 2024
Cited alongside, same era.
Steering Language Model Refusal with Sparse Autoencoders, November 2024
O’Brien, K., Majercak, D., Fernandes, X., Edgar, R., Chen, J., Nori, H., Carignan, D., Horvitz, E., and Poursabzi-Sangde, F · 2024
Later among the works it cites.
Automatically Interpreting Millions of Features in Large Language Models, October 2024
Paulo, G., Mallen, A., Juang, C., and Belrose, N · 2024
Later among the works it cites.
Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders
Rajamanoharan, S., Conmy, A., Smith, L., Lieberum, T., Varma, V., Kramar, J., Shah, R., and Nanda, N · 2024
Later among the works it cites.
Decomposing and editing predictions by modeling model computation, 2024
Shah, H., Ilyas, A., and Madry, A · 2024
Later among the works it cites.
Transformers use causal world models in maze-solving tasks, 2024
Spies, A. F., Edwards, W., Ivanitskiy, M. I., Skapars, A., Räuker, T., Inoue, K., Russo, A., and Shanahan, M · 2024
Later among the works it cites.
Attribution Patching Outperforms Automated Circuit Discovery
Syed, A., Rager, C., and Conmy, A · 2024
Later among the works it cites.
Predicting Future Activations, January 2024a
Templeton, A., Batson, J., Jermyn, A., and Olah, C · 2024
Later among the works it cites.
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet, May 2024b
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Tamkin, A., Durmus, E., Hume, T., Mosconi, F., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Later among the works it cites.
Brinkmann, J., Wendler, C., Bartelt, C., and Mueller, A · 2025
Closest in time.
Sparse autoencoders can interpret randomly initialized transformers, 2025
Heap, T., Lawson, T., Farnik, L., and Aitchison, L · 2025
Closest in time.