Fetching the paper…
Reading the bibliography…
For large language models (LLMs), sparse autoencoders (SAEs) have been shown to decompose intermediate representations that often are not interpretable directly into sparse sums of interpretable features, facilitating better control and subsequent analysis.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
Extensions of lipschitz maps into banach spaces
William B Johnson, Joram Lindenstrauss, and Gideon Schechtman · 1986
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space, 2013
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Describing textures in the wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, , and A. Vedaldi · 2014
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin · 2018
Earlier work this paper cites.
Gan dissection: Visualizing and understanding generative adversarial networks
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Bolei Zhou, Joshua B Tenenbaum, William T Freeman, and Antonio Torralba · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown · 2020
Earlier work this paper cites.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Earlier work this paper cites.
An overview of early vision in inceptionv1
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
GLU variants improve transformer
Noam Shazeer · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis, 2021
Prafulla Dhariwal and Alex Nichol · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Earlier work this paper cites.
Zeyu Yun, Yubei Chen, Bruno A Olshausen, and Yann LeCun · 2021
Earlier work this paper cites.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al · 2022
Earlier work this paper cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al · 2022
Earlier work this paper cites.
An image is worth multiple words: Multi-attribute inversion for constrained text-to-image synthesis
Aishwarya Agarwal, Srikrishna Karanam, Tripti Shukla, and Balaji Vasan Srinivasan · 2023
Cited alongside, same era.
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al · 2023
Cited alongside, same era.
Sega: Instructing text-to-image models using semantic guidance
Manuel Brack, Felix Friedrich, Dominik Hintersdorf, Lukas Struppek, Patrick Schramowski, and Kristian Kersting · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, and Adam Jermyn et al · 2023
Cited alongside, same era.
Reproducible scaling laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev · 2023
Cited alongside, same era.
Open source automated interpretability for sparse autoencoder features, 2024
Juang Caden, Paulo Gonçalo, Drori Jacob, and Belrose Nora · 2024
Closest in time.
Selfie: Self-interpretation of large language model embeddings
Haozhe Chen, Carl Vondrick, and Chengzhi Mao · 2024
Closest in time.
Noiseclr: A contrastive learning approach for unsupervised discovery of interpretable directions in diffusion models
Yusuf Dalva and Pinar Yanardag · 2024
Closest in time.
Interpreting and steering features in images, 2024
Gytis Daujotas · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Cited alongside, same era.
Concept sliders: Lora adaptors for precise control in diffusion models, 2023
Rohit Gandikota, Joanna Materzynska, Tingrui Zhou, Antonio Torralba, and David Bau · 2023
Cited alongside, same era.
Concept bottleneck generative models
Aya Abdelsalam Ismail, Julius Adebayo, Hector Corrada Bravo, Stephen Ra, and Kyunghyun Cho · 2023
Cited alongside, same era.
Direct inversion: Boosting diffusion-based editing with 3 lines of code, 2023
Xuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu, and Qiang Xu · 2023
Cited alongside, same era.
Diffusion models already have a semantic latent space, 2023
Mingi Kwon, Jaeseok Jeong, and Youngjung Uh · 2023
Cited alongside, same era.
Neuronpedia: Interactive reference and tooling for analyzing neural networks, 2023
Johnny Lin and Joseph Bloom · 2023
Cited alongside, same era.
Understanding the latent space of diffusion models through the lens of riemannian geometry
Yong-Hyun Park, Mingi Kwon, Jaewoong Choi, Junghyo Jo, and Youngjung Uh · 2023
Cited alongside, same era.
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva · 2024
Closest in time.
Self-discovering interpretable diffusion latent directions for responsible text-to-image generation
Hang Li, Chengzhi Shen, Philip Torr, Volker Tresp, and Jindong Gu · 2024
Closest in time.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2
Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca Dragan, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Sparse crosscoders for cross-layer features and model diffing
Jack Lindsey, Adly Templeton, Jonathan Marcus, Thomas Conerly, Joshua Batson, and Christopher Olah · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Hello gpt-4o, 2024
OpenAI · 2024
Closest in time.
Steering llama 2 via contrastive activation addition, 2024
Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner · 2024
Closest in time.
A practical review of mechanistic interpretability for transformer-based language models
Daking Rai, Yilun Zhou, Shi Feng, Abulhair Saparov, and Ziyu Yao · 2024
Closest in time.
Grounded sam: Assembling open-world models for diverse visual tasks, 2024
Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, Zhaoyang Zeng, Hao Zhang, Feng Li, Jie Yang, Hongyang Li, Qing Jiang, and Lei Zhang · 2024
Closest in time.
Laion coco: 600m synthetic captions from laion2b-en, 2022b
Christoph Schuhmann, Andreas Köpf, Richard Vencu, Theo Coombes, Romain Beaumont, and Benjamin Trom · 2024
Closest in time.
Interim research report: Taking features out of superposition with sparse autoencoders, 2022
Lee Sharkey, Dan Braun, and beren · 2024
Closest in time.
Advanced style transfer with the mad scientist node
Matteo Spinelli · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet, 2024
Adly Templeton and Tom Conerly et al · 2024
Closest in time.
Saeuron: Interpretable concept unlearning in diffusion models with sparse autoencoders, 2025
Bartosz Cywiński and Kamil Deja · 2025
Closest in time.
Learning on model weights using tree experts
Eliahu Horwitz, Bar Cavia, Jonathan Kahana, and Yedid Hoshen · 2025
Closest in time.