Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, et al · 2023
Later among the works it cites.
Sparse autoencoder library
Alan Cooney · 2023
Later among the works it cites.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs Smith, Robert Huben, and Lee Sharkey · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
The linear representation hypothesis and the geometry of large language models
Original
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Later among the works it cites.
Taking features out of superposition with sparse autoencoders, 2023
Lee Sharkey, Dan Braun, and Beren Millidge · 2023
Later among the works it cites.
Sae lens
Joseph Bloom and David Chanin · 2024
Closest in time.
A primer on the inner workings of transformer-based language models
Original
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R Costa-jussà · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Original
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Closest in time.
Tanh penalty in dictionary learning, 2024
Adam Jermyn, Adly Templeton, Joshua Batson, and Trenton Bricken · 2024
Closest in time.
Emergent world models and latent variable estimation in chess-playing language models, 2024
Adam Karvonen · 2024
Closest in time.
Interpreting attention layer outputs with sparse autoencoders
Connor Kissane, Robert Krzyzanowski, Joseph Isaac Bloom, Arthur Conmy, and Neel Nanda · 2024
Closest in time.
lichess.org open database, 2024
Lichess · 2024
Closest in time.
Towards principled evaluations of sparse autoencoders for interpretability and control
Original
Aleksandar Makelov, George Lange, and Neel Nanda · 2024
Closest in time.
dictionary_learning
Samuel Marks and Aaron Mueller · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Original
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Improving dictionary learning with gated sparse autoencoders
Original
Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Improving sparse autoencoders by square-rooting l1 and removing lowest activation features, 2024
Logan Riggs and Jannik Brinkmann · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.
Addressing feature suppression in sparse autoencoders, 2024
Benjamin Wright and Lee Sharkey · 2024
Closest in time.