Fetching the paper…
Reading the bibliography…
Sparse autoencoders (SAEs) are a promising approach to interpreting the internal representations of transformer language models.
An Information-Maximization Approach to Blind Separation and Blind Deconvolution
Anthony J. Bell and Terrence J. Sejnowski · 1995
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A. Olshausen and David J. Field · 1996
Earlier work this paper cites.
The “independent components” of natural scenes are edge filters
Anthony J. Bell and Terrence J. Sejnowski · 1997
Earlier work this paper cites.
Independent component analysis: algorithms and applications
A. Hyvärinen and E. Oja · 2000
Earlier work this paper cites.
Efficient sparse coding algorithms
Honglak Lee, Alexis Battle, Rajat Raina, and Andrew Ng · 2006
Earlier work this paper cites.
ICA with Reconstruction Cost for Efficient Overcomplete Feature Learning
Quoc Le, Alexandre Karpenko, Jiquan Ngiam, and Andrew Ng · 2011
Earlier work this paper cites.
Sparse autoencoder, 2011
Andrew Ng · 2011
Earlier work this paper cites.
k-Sparse Autoencoders, March 2014
Alireza Makhzani and Brendan Frey · 2014
Earlier work this paper cites.
Zero-bias autoencoders and the benefits of co-adapting features, April 2015
Kishore Konda, Roland Memisevic, and David Krueger · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization, January 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Residual Connections Encourage Iterative Inference, March 2018
Stanisław Jastrzębski, Devansh Arpit, Nicolas Ballas, Vikas Verma, Tong Che, and Yoshua Bengio · 2018
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners, 2019
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling, December 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Earlier work this paper cites.
Interpreting GPT: the logit lens, August 2020
nostalgebraist · 2020
Earlier work this paper cites.
Zoom In: An Introduction to Circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
A Mathematical Framework for Transformer Circuits, 2021
Nelson Elhage, Neel Nanda, Catherine Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, and T. Conerly · 2021
Earlier work this paper cites.
Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors
Zeyu Yun, Yubei Chen, Bruno Olshausen, and Yann LeCun · 2021
Earlier work this paper cites.
Toy Models of Superposition, September 2022
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Earlier work this paper cites.
Taking features out of superposition with sparse autoencoders, December 2022
Lee Sharkey, Dan Braun, and Beren Millidge · 2022
Cited alongside, same era.
Extracting Latent Steering Vectors from Pretrained Language Models
Nishant Subramani, Nivedita Suresh, and Matthew Peters · 2022
Cited alongside, same era.
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Cited alongside, same era.
High-Dimensional Data Analysis with Low-Dimensional Models: Principles, Computation, and Applications
John Wright and Yi Ma · 2022
Cited alongside, same era.
Eliciting Latent Predictions from Transformers with the Tuned Lens, November 2023
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
A Primer on the Inner Workings of Transformer-based Language Models, May 2024
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R. Costa-jussà · 2024
Closest in time.
Scaling and evaluating sparse autoencoders, June 2024
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Closest in time.
How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability
Jorge García-Carrasco, Alejandro Maté, and Juan Carlos Trujillo · 2024
Closest in time.
The Llama 3 Herd of Models, November 2024
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal · 2023
Cited alongside, same era.
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning, 2023
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, and Amanda Askell · 2023
Cited alongside, same era.
Towards Automated Circuit Discovery for Mechanistic Interpretability
Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Cited alongside, same era.
Sparse Autoencoders Find Highly Interpretable Features in Language Models, October 2023
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Cited alongside, same era.
openai/sparse_autoencoder, December 2023
Leo Gao, Tom Dupré la Tour, and Jeffrey Wu · 2023
Cited alongside, same era.
Residual stream norms grow exponentially over the forward pass, May 2023
Stefan Heimersheim and Alex Turner · 2023
Cited alongside, same era.
The Linear Representation Hypothesis and the Geometry of Large Language Models, November 2023
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Cited alongside, same era.
Zhengfu He, Xuyang Ge, Qiong Tang, Tianxiang Sun, Qinyuan Cheng, and Xipeng Qiu · 2024
Closest in time.
Linearity of Relation Decoding in Transformer Language Models, February 2024
Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau · 2024
Closest in time.
Jing Huang, Zhengxuan Wu, Christopher Potts, Mor Geva, and Atticus Geiger · 2024
Closest in time.
Interpreting Attention Layer Outputs with Sparse Autoencoders
Connor Kissane, Robert Krzyzanowski, Joseph Isaac Bloom, Arthur Conmy, and Neel Nanda · 2024
Closest in time.
The Remarkable Robustness of LLMs: Stages of Inference?, June 2024
Vedang Lad, Wes Gurnee, and Max Tegmark · 2024
Closest in time.
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2, August 2024
Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca Dragan, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Sparse Autoencoders Match Supervised Features for Model Steering on the IOI Task
Aleksandar Makelov · 2024
Closest in time.
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
The Next Five Hurdles, July 2024
Chris Olah · 2024
Closest in time.
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models, May 2024
Charles O’Neill and Thang Bui · 2024
Closest in time.
Disentangling Dense Embeddings with Sparse Autoencoders, August 2024
Charles O’Neill, Christine Ye, Kartheik Iyer, and John F. Wu · 2024
Closest in time.
Gemma 2: Improving Open Language Models at a Practical Size, October 2024
Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, et al · 2024
Closest in time.
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet, May 2024
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L. Turner, Callum McDougall, Monte MacDiarmid, Alex Tamkin, Esin Durmus, Tristan Hume, Francesco Mosconi, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.
Relational Composition in Neural Networks: A Survey and Call to Action
Martin Wattenberg and Fernanda Viégas · 2024
Closest in time.