Fetching the paper…
Reading the bibliography…
Sparse autoencoders (SAEs) are an unsupervised method for learning a sparse decomposition of a neural network's latent representations into seemingly interpretable features.
Analysing mathematical reasoning abilities of neural models, 2019
D. Saxton, E. Grefenstette, F. Hill, and P. Kohli · 1904
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 1909
Earlier work this paper cites.
Root mean square layer normalization, 2019
B. Zhang and R. Sennrich · 1910
Earlier work this paper cites.
Language models are few-shot learners, 2020
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2005
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
P. Koehn · 2005
Earlier work this paper cites.
Efficient estimation of word representations in vector space, 2013
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification, 2015
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
D. P. Kingma and J. Ba · 2017
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2020
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter · 2020
Earlier work this paper cites.
Branch specialization
C. Voss, G. Goh, N. Cammarata, M. Petrov, L. Schubert, and C. Olah · 2020
Earlier work this paper cites.
Interpretability
C. Olah · 2021
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Y. Belinkov · 2022
Earlier work this paper cites.
Toy models of superposition
N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, R. Grosse, S. McCandlish, J. Kaplan, D. Amodei, M. Wattenberg, and C. Olah · 2022
Earlier work this paper cites.
A transparency and interpretability tech tree
E. Hubinger · 2022
Earlier work this paper cites.
A longlist of theories of impact for interpretability
N. Nanda · 2022
Earlier work this paper cites.
Transformerlens
N. Nanda and J. Bloom · 2022
Earlier work this paper cites.
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, S. Johnston, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2022
Earlier work this paper cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2022
Cited alongside, same era.
Adversarial training for high-stakes reliability
D. Ziegler, S. Nix, L. Chan, T. Bauman, P. Schmidt-Nielsen, T. Lin, A. Scherlis, N. Nabeshima, B. Weinstein-Raun, D. de Haas, B. Shlegeris, and N. Thomas · 2022
Cited alongside, same era.
Language models can explain neurons in language models
S. Bills, N. Cammarata, D. Mossing, H. Tillman, L. Gao, G. Goh, I. Sutskever, J. Leike, J. Wu, and W. Saunders · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, Z. Hatfield-Dodds, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability, 2023
Update on how we train SAEs
T. Conerly, A. Templeton, T. Bricken, J. Marcus, and T. Henighan · 2024
Closest in time.
Activation steering with SAEs
A. Conmy and N. Nanda · 2024
Closest in time.
Circuits Updates - June 2024: Comparing TopK and Gated SAEs to Standard SAEs
H. Cunningham and T. Conerly · 2024
Closest in time.
Transcoders find interpretable llm feature circuits, 2024
J. Dunefsky, P. Chlenski, and N. Nanda · 2024
Closest in time.
Not all language model features are linear, 2024
J. Engels, I. Liao, E. J. Michaud, W. Gurnee, and M. Tegmark · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Conmy, A. N. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models, 2023
H. Cunningham, A. Ewart, L. Riggs, R. Huben, and L. Sharkey · 2023
Cited alongside, same era.
Finding neurons in a haystack: Case studies with sparse probing
W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas · 2023
Cited alongside, same era.
M. Hanna, O. Liu, and A. Variengien · 2023
Cited alongside, same era.
In-context learning creates task vectors
R. Hendel, M. Geva, and A. Globerson · 2023
Cited alongside, same era.
Rigorously assessing natural language explanations of neurons, 2023
J. Huang, A. Geiger, K. D’Oosterlinck, Z. Wu, and C. Potts · 2023
Cited alongside, same era.
Attention head superposition, May 2023
A. Jermyn, C. Olah, and T. Henighan · 2023
Cited alongside, same era.
Inference-time intervention: Eliciting truthful answers from a language model
K. Li, O. Patel, F. Viégas, H. Pfister, and M. Wattenberg · 2023
Cited alongside, same era.
L. Gao, T. D. la Tour, H. Tillman, G. Goh, R. Troll, A. Radford, I. Sutskever, J. Leike, and J. Wu · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Gemini Team · 2024
Closest in time.
Universal neurons in gpt2 language models, 2024
W. Gurnee, T. Horsley, Z. C. Guo, T. R. Kheirkhah, Q. Sun, W. Hathaway, N. Nanda, and D. Bertsimas · 2024
Closest in time.
What makes and breaks safety fine-tuning? a mechanistic study, 2024
S. Jain, E. S. Lubana, K. Oksuz, T. Joy, P. H. S. Torr, A. Sanyal, and P. K. Dokania · 2024
Closest in time.
Measuring progress in dictionary learning for language model interpretability with board game models
A. Karvonen, B. Wright, C. Rager, R. Angell, J. Brinkmann, L. R. Smith, C. M. Verdun, D. Bau, and S. Marks · 2024
Closest in time.
Towards principled evaluations of sparse autoencoders for interpretability and control
A. Makelov, G. Lange, and N. Nanda · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models, 2024
S. Marks, C. Rager, E. J. Michaud, Y. Belinkov, D. Bau, and A. Mueller · 2024
Closest in time.
Open Problem: Attribution Dictionary Learning
C. Olah, A. Templeton, T. Bricken, and A. Jermyn · 2024
Closest in time.
Replacing sae encoders with inference-time optimisation, 2024
L. Smith · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
A. Templeton, T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones, H. Cunningham, N. L. Turner, C. McDougall, M. MacDiarmid, C. D. Freeman, T. R. Sumers, E. Rees, J. Batson, A. Jermyn, S. Carter, C. Olah, and T. Henighan · 2024
Closest in time.
Function vectors in large language models
E. Todd, M. Li, A. S. Sharma, A. Mueller, B. C. Wallace, and D. Bau · 2024
Closest in time.
Activation addition: Steering language models without optimization, 2024
A. M. Turner, L. Thiergart, G. Leech, D. Udell, J. J. Vazquez, U. Mini, and M. MacDiarmid · 2024
Closest in time.
Relational composition in neural networks: A gentle survey and call to action
M. Wattenberg and F. Viégas · 2024
Closest in time.