Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors, 2021
Original
Zeyu Yun, Yubei Chen, Bruno A Olshausen, and Yann LeCun · 2021
Cited alongside, same era.
Toy models of superposition, 2022
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Cited alongside, same era.
Attribution patching: Activation patching at industrial scale
Neel Nanda · 2022
Cited alongside, same era.
Transformerlens
Neel Nanda and Joseph Bloom · 2022
Cited alongside, same era.
Taking features out of superposition with sparse autoencoders, Dec 2022
Lee Sharkey, Dan Braun, and Beren Millidge · 2022
Cited alongside, same era.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Original
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Cited alongside, same era.
Tinystories: How small can language models be and still speak coherent english?
Original
Ronen Eldan and Yuanzhi Li · 2023
Cited alongside, same era.
Analyzing neural networks with dictionary learning, 2023
Johnny Lin and Joseph Bloom · 2023
Cited alongside, same era.
Codebook features: Sparse and discrete interpretability for neural networks
Alex Tamkin, Mohammad Taufeeque, and Noah D. Goodman · 2023
Cited alongside, same era.