Fetching the paper…
Reading the bibliography…
Sparse autoencoders provide a promising unsupervised approach for extracting interpretable features from a language model by reconstructing activations from a sparse bottleneck layer.
Ben Zhou, Daniel Khashabi, Qiang Ning, and Dan Roth · 1909
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
Matching pursuits with time-frequency dictionaries
Stéphane G Mallat and Zhifeng Zhang · 1993
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A Olshausen and David J Field · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
The jpeg 2000 still image compression standard
Athanassios Skodras, Charilaos Christopoulos, and Touradj Ebrahimi · 2001
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn · 2005
Earlier work this paper cites.
K-SVD: An algorithm for designing overcomplete dictionaries for sparse representation
Michal Aharon, Michael Elad, and Alfred Bruckstein · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Sparse deep belief net model for visual area v2
Honglak Lee, Chaitanya Ekanadham, and Andrew Ng · 2007
Earlier work this paper cites.
Relaxed lasso
Nicolai Meinshausen · 2007
Earlier work this paper cites.
Coherence analysis of iterative thresholding algorithms
Arian Maleki · 2009
Earlier work this paper cites.
How to explain individual classification decisions
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert Müller · 2010
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Quoc V Le, Marc’Aurelio Ranzato, Rajat Monga, Matthieu Devin, Kai Chen, Greg S Corrado, Jeff Dean, and Andrew Y Ng · 2013
Earlier work this paper cites.
Alireza Makhzani and Brendan Frey · 2013
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
Julian McAuley and Jure Leskovec · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Zero-bias autoencoders and the benefits of co-adapting features
Kishore Konda, Roland Memisevic, and David Krueger · 2014
Earlier work this paper cites.
Sparse modeling for image and vision processing
Julien Mairal, Francis Bach, Jean Ponce, et al · 2014
Earlier work this paper cites.
Toxic comment classification challenge
Cjadams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, Nithum, and Will Cukierski · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F. Liu, Matt Gardner, Gabor Angeli, Rik Koncel-Kedziorski, Emily Bender, Kyle Richardson, Peter Clark, and Nate Kushman · 2017
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction
Leland McInnes, John Healy, and James Melville · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Cited alongside, same era.
JumpReLU: A retrofit defense strategy for adversarial attacks
N Benjamin Erichson, Zhewei Yao, and Michael W Mahoney · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Neuron to graph: Interpreting language model neurons at scale
Alex Foote, Neel Nanda, Esben Kran, Ioannis Konstas, Shay Cohen, and Fazl Barez · 2023
Later among the works it cites.
Finding neurons in a haystack: Case studies with sparse probing
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas · 2023
Later among the works it cites.
Some open-source dictionaries and dictionary learning infrastructure
Sam Marks · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Socialiqa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Quartz: An open-domain dataset of qualitative relationship questions
Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Cited alongside, same era.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Later among the works it cites.
Pytorch fsdp: experiences on scaling fully sharded data parallel
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al · 2023
Later among the works it cites.
Open source sparse autoencoders for all residual stream layers of gpt2-small
Joseph Bloom · 2024
Closest in time.
Identifying functionally important features with end-to-end sparse dictionary learning
Dan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, and Lee Sharkey · 2024
Closest in time.
Update on how we train saes
Tom Conerly, Adly Templeton, Trenton Bricken, Jonathan Marcus, and Tom Henighan · 2024
Closest in time.
Decoding the thought vector, 2016
Gabriel Goh · 2024
Closest in time.
Ag’s corpus of news articles
Antonio Gulli · 2024
Closest in time.
Ghost grads: An improvement on resampling
Adam Jermyn and Adly Templeton · 2024
Closest in time.
Scaling laws for dictionary learning
Jack Lindsey, Tom Conerly, Adly Templeton, Jonathan Marcus, and Tom Henighan · 2024
Closest in time.
Towards principled evaluations of sparse autoencoders for interpretability and control, 2024
Aleksandar Makelov, George Lange, and Neel Nanda · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Transformer debugger
Dan Mossing, Steven Bills, Henk Tillman, Tom Dupré la Tour, Nick Cammarata, Leo Gao, Joshua Achiam, Catherine Yeh, Jan Leike, Jeff Wu, and William Saunders · 2024
Closest in time.
Progress update #1 from the gdm mech interp team: Full update
Neel Nanda, Arthur Conmy, Lewis Smith, Senthooran Rajamanoharan, Tom Lieberum, János Kramár, and Vikrant Varma · 2024
Closest in time.
Open problem: Attribution dictionary learning
Chris Olah, Adly Templeton, Trenton Bricken, and Adam Jermyn · 2024
Closest in time.
Gpt-2 output dataset
OpenAI · 2024
Closest in time.
Improving dictionary learning with gated sparse autoencoders
Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Massive activations in large language models
Mingjie Sun, Xinlei Chen, J Zico Kolter, and Zhuang Liu · 2024
Closest in time.
ProLU: A nonlinearity for sparse autoencoders
Glen Taggart · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.
Addressing feature suppression in SAEs
Benjamin Wright and Lee Sharkey · 2024
Closest in time.