Fetching the paper…
Reading the bibliography…
Sparse autoencoders (SAEs) are a promising unsupervised approach for identifying causally relevant and interpretable linear features in a language model's (LM) activations.
Understanding straight-through estimator in training activation quantized neural nets, 2019
P. Yin, J. Lyu, S. Zhang, S. Osher, Y. Qi, and J. Xin · 1903
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library, 2019
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 1912
Earlier work this paper cites.
The perceptron: A probabilistic model for information storage and organization in the brain
F. Rosenblatt · 1958
Earlier work this paper cites.
On Estimation of a Probability Density Function and Mode
E. Parzen · 1962
Earlier work this paper cites.
Matching pursuits with time-frequency dictionaries
S. Mallat and Z. Zhang · 1993
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
B. A. Olshausen and D. J. Field · 1997
Earlier work this paper cites.
All of statistics : a concise course in statistical inference
L. Wasserman · 2010
Earlier work this paper cites.
Sparse autoencoder
A. Ng · 2011
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation, 2013
Y. Bengio, N. Léonard, and A. Courville · 2013
Earlier work this paper cites.
Zero-bias autoencoders and the benefits of co-adapting features, 2015
K. Konda, R. Memisevic, and D. Krueger · 2015
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations, 2016
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio · 2016
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
D. P. Kingma and J. Ba · 2017
Earlier work this paper cites.
A. Athalye, N. Carlini, and D. Wagner · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
Learning sparse neural networks through l 0 l_{0} regularization
C. Louizos, M. Welling, and D. P. Kingma · 2018
Earlier work this paper cites.
Jumprelu: A retrofit defense strategy for adversarial attacks, 2019
N. B. Erichson, Z. Yao, and M. W. Mahoney · 2019
Cited alongside, same era.
Zoom in: An introduction to circuits
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2021
Cited alongside, same era.
Tpu-knn: K nearest neighbor search at peak flop/s, 2022
F. Chern, B. Hechtman, A. Davis, R. Guo, D. Majnemer, and S. Kumar · 2022
Cited alongside, same era.
[interim research report] taking features out of superposition with sparse autoencoders, 2022
L. Sharkey, D. Braun, and B. Millidge · 2022
Cited alongside, same era.
Update on how we train SAEs
T. Conerly, A. Templeton, T. Bricken, J. Marcus, and T. Henighan · 2024
Closest in time.
Activation steering with SAEs
A. Conmy and N. Nanda · 2024
Closest in time.
Circuits Updates - June 2024: Comparing TopK and Gated SAEs to Standard SAEs
H. Cunningham and T. Conerly · 2024
Closest in time.
Transcoders find interpretable llm feature circuits, 2024
J. Dunefsky, P. Chlenski, and N. Nanda · 2024
Closest in time.
Scaling and evaluating sparse autoencoders, 2024
L. Gao, T. D. la Tour, H. Tillman, G. Goh, R. Troll, A. Radford, I. Sutskever, J. Leike, and J. Wu · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pythia: A suite for analyzing large language models across training and scaling
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, et al · 2023
Cited alongside, same era.
Language models can explain neurons in language models
S. Bills, N. Cammarata, D. Mossing, H. Tillman, L. Gao, G. Goh, I. Sutskever, J. Leike, J. Wu, and W. Saunders · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, Z. Hatfield-Dodds, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability, 2023
A. Conmy, A. N. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models, 2023
H. Cunningham, A. Ewart, L. Riggs, R. Huben, and L. Sharkey · 2023
Cited alongside, same era.
Extremely simple activation shaping for out-of-distribution detection, 2023
A. Djurisic, N. Bozanic, A. Ashok, and R. Liu · 2023
Cited alongside, same era.
Finding neurons in a haystack: Case studies with sparse probing, 2023
W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas · 2023
Cited alongside, same era.
Gemini Team · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size, 2024
Gemma Team · 2024
Closest in time.
Interpreting sae features with gemini ultra
T. Lieberum · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models, 2024
S. Marks, C. Rager, E. J. Michaud, Y. Belinkov, D. Bau, and A. Mueller · 2024
Closest in time.
Open Problem: Attribution Dictionary Learning
C. Olah, A. Templeton, T. Bricken, and A. Jermyn · 2024
Closest in time.
Improving dictionary learning with gated sparse autoencoders, 2024
S. Rajamanoharan, A. Conmy, L. Smith, T. Lieberum, V. Varma, J. Kramár, R. Shah, and N. Nanda · 2024
Closest in time.
Prolu: A nonlinearity for sparse autoencoders
G. M. Taggart · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
A. Templeton, T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones, H. Cunningham, N. L. Turner, C. McDougall, M. MacDiarmid, C. D. Freeman, T. R. Sumers, E. Rees, J. Batson, A. Jermyn, S. Carter, C. Olah, and T. Henighan · 2024
Closest in time.
Addressing feature suppression in saes
B. Wright and L. Sharkey · 2024
Closest in time.