Fetching the paper…
Reading the bibliography…
Sparse Autoencoders (SAEs) have emerged as a useful tool for interpreting the internal representations of neural networks.
A mathematical theory of communication
C. E. Shannon · 1948
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
An introduction to coding theory and the two-part minimum description length principle
T. C. Lee · 2001
Earlier work this paper cites.
Information theory, inference and learning algorithms
D. J. MacKay · 2003
Earlier work this paper cites.
The minimum description length principle
P. D. Grünwald · 2007
Earlier work this paper cites.
Variable-length codes for data compression
D. Salomon · 2007
Earlier work this paper cites.
Minimum description length penalization for group and multi-task sparse learning
P. S. Dhillon, D. Foster, and L. H. Ungar · 2011
Earlier work this paper cites.
An mdl framework for sparse coding and dictionary learning
I. Ramirez and G. Sapiro · 2012
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Q. V. Le · 2013
Cited alongside, same era.
A. Makhzani and B. Frey · 2013
Cited alongside, same era.
Rethinking lossy compression: The rate-distortion-perception tradeoff
Y. Blau and T. Michaeli · 2019
Cited alongside, same era.
Taking features out of superposition with sparse autoencoders
L. Sharkey, D. Braun, and B. Millidge · 2022
Cited alongside, same era.
Language models can explain neurons in language models
S. Bills, N. Cammarata, D. Mossing, H. Tillman, L. Gao, G. Goh, I. Sutskever, J. Leike, J. Wu, and W. Saunders · 2023
Cited alongside, same era.
Features as the simplest factorization
Open source sparse autoencoders for all residual stream layers of gpt2-small
J. Bloom · 2024
Closest in time.
Batchtopk: A simple improvement for topk-saes
B. Bussmann, P. Leask, and N. Nanda · 2024
Closest in time.
Compact proofs of model performance via mechanistic interpretability
L. Chan, R. Agrawal, A. Garriga-Alonso, and J. Gross · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
L. Gao, T. D. la Tour, H. Tillman, G. Goh, R. Troll, A. Radford, I. Sutskever, J. Leike, and J. Wu · 2024
Closest in time.
Sparse autoencoders find highly interpretable features in language models
R. Huben, H. Cunningham, L. R. Smith, A. Ewart, and L. Sharkey · 2024
Closest in time.
Circuits updates - july 2024, linear representations
C. Olah · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Bricken, J. Batson, A. Templeton, A. Jermyn, T. Henighan, and C. Olah · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, Z. Hatfield-Dodds, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah · 2023
Cited alongside, same era.
Closest in time.
Sparsify: A mechanistic interpretability research agenda
L. Sharkey · 2024
Closest in time.