Fetching the paper…
Reading the bibliography…
Sparse autoencoders (SAEs) are a promising technique for decomposing language model activations into interpretable linear features.
Robust principal component analysis?
Emmanuel J Candès, Xiaodong Li, Yi Ma, and John Wright · 2011
Earlier work this paper cites.
A Mathematical Introduction to Compressive Sensing
Simon Foucart and Holger Rauhut · 2013
Earlier work this paper cites.
Alireza Makhzani and Brendan Frey · 2013
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain · 2016
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Towards improving discriminative reconstruction via simultaneous dense and sparse coding
Abiy Tasissa, Emmanouil Theodosis, Bahareh Tolooshams, and Demba Ba · 2020
Earlier work this paper cites.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Earlier work this paper cites.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Earlier work this paper cites.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Earlier work this paper cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Earlier work this paper cites.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Earlier work this paper cites.
Interpretability dreams
Chris Olah · 2023
Cited alongside, same era.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Cited alongside, same era.
Llama 3 model card, 2024
AI@Meta · 2024
Cited alongside, same era.
Examining language model performance with reconstructed activations using sparse autoencoders
Evan Anders and Joseph Bloom · 2024
Cited alongside, same era.
Circuits updates april 2024, 2024
Transformer Circuits Team Anthropic · 2024
Cited alongside, same era.
Open source sparse autoencoders for all residual stream layers of gpt2 small
Joseph Bloom · 2024
Cited alongside, same era.
Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders
Zhengfu He, Wentao Shu, Xuyang Ge, Lingjie Chen, Junxuan Wang, Yunhua Zhou, Frances Liu, Qipeng Guo, Xuanjing Huang, Zuxuan Wu, et al · 2024
Closest in time.
Activation plateaus & sensitive directions in gpt2
Stefan Heimersheim and Jake Mendel · 2024
Closest in time.
Open source automated interpretability for sparse autoencoder features
Caden Juang, Gonçalo Paulo, Jacob Drori, and Nora Belrose · 2024
Closest in time.
Measuring progress in dictionary learning for language model interpretability with board game models
Adam Karvonen, Benjamin Wright, Can Rager, Rico Angell, Jannik Brinkmann, Logan Smith, Claudio Mayrink Verdun, David Bau, and Samuel Marks · 2024
Closest in time.
The remarkable robustness of llms: Stages of inference?
Vedang Lad, Wes Gurnee, and Max Tegmark · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stitching saes of different sizes
Bart Bussmann, Patrick Leask, Joseph Bloom, Curt Tigges, and Neel Nanda · 2024
Cited alongside, same era.
Recurrent neural networks learn to store and generate sequences using non-linear representations
Róbert Csordás, Christopher Potts, Christopher D Manning, and Atticus Geiger · 2024
Cited alongside, same era.
Not all language model features are linear
Joshua Engels, Isaac Liao, Eric J Michaud, Wes Gurnee, and Max Tegmark · 2024
Cited alongside, same era.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Cited alongside, same era.
Evaluating synthetic activations composed of sae latents in gpt-2
Giorgi Giglemiani, Nora Petrova, Chatrik Singh Mangat, Jett Janiak, and Stefan Heimersheim · 2024
Cited alongside, same era.
Sae reconstruction errors are (empirically) pathological
Wes Gurnee · 2024
Cited alongside, same era.
Closest in time.
Investigating sensitive directions in gpt-2: An improved baseline and comparative analysis of saes
Daniel Lee and Stefan Heimersheim · 2024
Closest in time.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2
Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca Dragan, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Sae feature geometry is outside the superposition hypothesis
Jake Mendel · 2024
Closest in time.
Jumping ahead: Improving reconstruction fidelity with jumprelu sparse autoencoders
Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, János Kramár, and Neel Nanda · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.