Fetching the paper…
Reading the bibliography…
Sparse Autoencoders (SAEs) are widely used to interpret neural networks by identifying meaningful concepts from their representations.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A Olshausen and David J Field · 1996
Earlier work this paper cites.
Wavelet-like receptive fields emerge from a network that learns sparse codes for natural images
Bruno A Olshausen and David J Field · 1996
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A Olshausen and David J Field · 1997
Earlier work this paper cites.
Dictionary learning algorithms for sparse representation
Kenneth Kreutz-Delgado, Joseph F Murray, Bhaskar D Rao, Kjersti Engan, Te-Won Lee, and Terrence J Sejnowski · 2003
Earlier work this paper cites.
An iterative thresholding algorithm for linear inverse problems with a sparsity constraint
Ingrid Daubechies, Michel Defrise, and Christine De Mol · 2004
Earlier work this paper cites.
Learning a dictionary of shape-components in visual cortex: Comparison with neurons, humans and machines
Thomas Serre · 2006
Earlier work this paper cites.
Analysis versus synthesis in signal priors
Michael Elad, Peyman Milanfar, and Ron Rubinstein · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
Dictionary learning and sparse coding for unsupervised clustering
Pablo Sprechmann and Guillermo Sapiro · 2010
Earlier work this paper cites.
Learning fast approximations of sparse coding
Karol Gregor and Yann LeCun · 2010
Earlier work this paper cites.
Sparse autoencoder
Andrew Ng et al · 2011
Earlier work this paper cites.
Alireza Makhzani and Brendan Frey · 2013
Earlier work this paper cites.
Sparse overcomplete word vector representations
Manaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, and Noah Smith · 2015
Earlier work this paper cites.
On the computational intractability of exact and approximate dictionary learning
Andreas M. Tillmann · 2015
Earlier work this paper cites.
Deep networks for image super-resolution with sparse prior
Zhaowen Wang, Ding Liu, Jianchao Yang, Wei Han, and Thomas Huang · 2015
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Andre Martins and Ramon Astudillo · 2016
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher P Burgess, Xavier Glorot, Matthew M Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Earlier work this paper cites.
Spine: Sparse interpretable neural embeddings
Anant Subramanian, Danish Pruthi, Harsh Jhamtani, Taylor Berg-Kirkpatrick, and Eduard Hovy · 2018
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
Challenging common assumptions in the unsupervised learning of disentangled representations
Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem · 2019
Earlier work this paper cites.
Debugging tests for model explanations
Julius Adebayo, Michael Muelly, Ilaria Liccardi, and Been Kim · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Weakly-supervised disentanglement without compromises
Francesco Locatello, Ben Poole, Gunnar Ratsch, Bernhard Scholkopf, Olivier Bachem, and Michael Tschannen · 2020
Earlier work this paper cites.
The incomplete rosetta stone problem: Identifiability results for multi-view nonlinear ica
Luigi Gresele, Paul K Rubenstein, Arash Mehrjou, Francesco Locatello, and Bernhard Scholkopf · 2020
Earlier work this paper cites.
Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing
Vishal Monga, Yuelong Li, and Yonina C Eldar · 2021
Earlier work this paper cites.
Graph unrolling networks: Interpretable neural networks for graph signal denoising
Siheng Chen, Yonina C Eldar, and Lingxiao Zhao · 2021
Earlier work this paper cites.
Self-supervised learning with data augmentations provably isolates content from style
Julius Von Kugelgen, Yash Sharma, Luigi Gresele, Wieland Brendel, Bernhard Scholkopf, Michel Besserve, and Francesco Locatello · 2021
Cited alongside, same era.
Interpretable machine learning: Fundamental principles and 10 grand challenges
Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong · 2022
Cited alongside, same era.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Cited alongside, same era.
Interpretable neural network via algorithm unrolling for mechanical fault diagnosis
Botao An, Shibin Wang, Zhibin Zhao, Fuhua Qin, Ruqiang Yan, and Xuefeng Chen · 2022
Cited alongside, same era.
Sparsity-constrained optimal transport
Tianlin Liu, Joan Puigcerver, and Mathieu Blondel · 2022
Recurrent neural networks learn to store and generate sequences using non-linear representations, 2024
Róbert Csordás, Christopher Potts, Christopher D. Manning, and Atticus Geiger · 2024
Later among the works it cites.
Not all language model features are linear, 2024
Joshua Engels, Eric J. Michaud, Isaac Liao, Wes Gurnee, and Max Tegmark · 2024
Later among the works it cites.
A percolation model of emergence: Analyzing transformers trained on a formal language
Ekdeep Singh Lubana, Kyogo Kawaguchi, Robert P Dick, and Hidenori Tanaka · 2024
Later among the works it cites.
Interpretable deep learning for deconvolutional analysis of neural signals
Bahareh Tolooshams, Sara Matias, Hao Wu, Simona Temereanca, Naoshige Uchida, Venkatesh N Murthy, Paul Masset, and Demba Ba · 2024
Later among the works it cites.
Improving dictionary learning with gated sparse autoencoders
Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah, and Neel Nanda · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Cited alongside, same era.
Craft: Concept recursive activation factorization for explainability
Thomas Fel, Agustin Picard, Louis Bethune, Thibaut Boissin, David Vigouroux, Julien Colin, Rémi Cadène, and Thomas Serre · 2023
Cited alongside, same era.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Cited alongside, same era.
K-deep simplex: Manifold learning via local dictionaries
Abiy Tasissa, Pranay Tankala, James M Murphy, and Demba Ba · 2023
Cited alongside, same era.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Samyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P Dick, Hidenori Tanaka, Edward Grefenstette, Tim Rocktäschel, and David Scott Krueger · 2023
Cited alongside, same era.
Physics of language models: Part 1, learning hierarchical language structures
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Cited alongside, same era.
Later among the works it cites.
Prolu: A nonlinearity for sparse autoencoders
Glen M. Taggart · 2024
Later among the works it cites.
Decomposing the dark matter of sparse autoencoders
Joshua Engels, Logan Riggs, and Max Tegmark · 2024
Later among the works it cites.
Saes are highly dataset dependent: A case study on the refusal direction
Connor Kissane, Robert Krzyzanowski, Neel Nanda, and Arthur Conmy · 2024
Later among the works it cites.
Interpretability as compression: Reconsidering sae explanations of neural activations with mdl-saes
Kola Ayonrinde, Michael T Pearce, and Lee Sharkey · 2024
Later among the works it cites.
Sparse crosscoders for cross-layer features and model diffing
Jack Lindsey, Adly Templeton, Jonathan Marcus, Thomas Conerly, Joshua Batson, and Christopher Olah · 2024
Later among the works it cites.
The geometry of categorical and hierarchical concepts in large language models
Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch · 2024
Later among the works it cites.
International ai safety report
Yoshua Bengio, Sören Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Philip Fox, Ben Garfinkel, Danielle Goldfarb, et al · 2025
Closest in time.
Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts, Susan Wei, Alexander Gietelink Oldenziel, George Wang, Liam Carroll, and Daniel Murfet · 2025
Closest in time.
Archetypal sae: Adaptive and stable dictionary learning for concept extraction in large vision models, 2025
Thomas Fel, Ekdeep Singh Lubana, Jacob S. Prince, Matthew Kowal, Victor Boutin, Isabel Papadimitriou, Binxu Wang, Martin Wattenberg, Demba Ba, and Talia Konkle · 2025
Closest in time.
Sparks of explainability: Recent advancements in explaining large vision models
Thomas Fel · 2025
Closest in time.
Universal sparse autoencoders: Interpretable cross-model concept alignment, 2025
Harrish Thasarathan, Julian Forsyth, Thomas Fel, Matthew Kowal, and Konstantinos Derpanis · 2025
Closest in time.
From mechanistic interpretability to mechanistic biology: Training, evaluating, and interpreting sparse autoencoders on protein language models
Etowah Adams, Liam Bai, Minji Lee, Yiyang Yu, and Mohammed AlQuraishi · 2025
Closest in time.
Interpreting and steering protein language models through sparse autoencoders
Edith Natalia Villegas Garcia and Alessio Ansuini · 2025
Closest in time.
Are sparse autoencoders useful? a case study in sparse probing
Subhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan, Max Tegmark, and Neel Nanda · 2025
Closest in time.
Sparse autoencoders trained on the same data learn different features
Gonçalo Paulo and Nora Belrose · 2025
Closest in time.
Truth is universal: Robust detection of lies in llms
Lennart Bürger, Fred A Hamprecht, and Boaz Nadler · 2025
Closest in time.
The hidden dimensions of llm alignment: A multi-dimensional safety analysis, 2025
Wenbo Pan, Zhichao Liu, Qiguang Chen, Xiangyang Zhou, Haining Yu, and Xiaohua Jia · 2025
Closest in time.
Transcoders find interpretable llm feature circuits
Jacob Dunefsky, Philippe Chlenski, and Neel Nanda · 2025
Closest in time.
Transcoders beat sparse autoencoders for interpretability
Gonçalo Paulo, Stepan Shabalin, and Nora Belrose · 2025
Closest in time.
Axbench: Steering llms? even simple baselines outperform sparse autoencoders
Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher D Manning, and Christopher Potts · 2025
Closest in time.
Sparse autoencoders do not find canonical units of analysis
Patrick Leask, Bart Bussmann, Michael Pearce, Joseph Bloom, Curt Tigges, Noura Al Moubayed, Lee Sharkey, and Neel Nanda · 2025
Closest in time.
Reft: Representation finetuning for language models
Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Dan Jurafsky, Christopher D Manning, and Christopher Potts · 2025
Closest in time.