Fetching the paper…
Reading the bibliography…
Sparse autoencoders (SAEs) have shown promise in extracting interpretable features from complex neural networks.
“Language Models are Few-Shot Learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 1901
Earlier work this paper cites.
“The vocabulary problem in human-system communication”
George Furnas, Thomas Landauer, Louis Gomez and Susan Dumais · 1987
Earlier work this paper cites.
“Indexing by latent semantic analysis”
Scott Deerwester, Susan Dumais, George Furnas, Thomas Landauer and Richard Harshman · 1990
Earlier work this paper cites.
“Sparse coding with an overcomplete basis set: A strategy employed by V1?”
Bruno Olshausen and David Field · 1997
Earlier work this paper cites.
“Modern information retrieval”
Ricardo Baeza-Yates and Berthier Ribeiro-Neto · 1999
Earlier work this paper cites.
“Mapping the backbone of science”
Kevin Boyack, Richard Klavans and Katy B"orner · 2005
Earlier work this paper cites.
“Compressed sensing”
David Donoho · 2006
Earlier work this paper cites.
“Introduction to information retrieval”
Christopher Manning, Prabhakar Raghavan and Hinrich Schütze · 2008
Earlier work this paper cites.
“Word representations: a simple and general method for semi-supervised learning”
Joseph Turian, Lev Ratinov and Yoshua Bengio · 2010
Earlier work this paper cites.
“Sparse autoencoder”
Andrew Ng · 2011
Earlier work this paper cites.
Alireza Makhzani and Brendan Frey · 2013
Earlier work this paper cites.
“Efficient estimation of word representations in vector space”
Tomas Mikolov, Kai Chen, Greg Corrado and Jeffrey Dean · 2013
Earlier work this paper cites.
“Distributed representations of words and phrases and their compositionality”
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado and Jeff Dean · 2013
Earlier work this paper cites.
“Atypical combinations and scientific impact”
Brian Uzzi, Satyam Mukherjee, Michael Stringer and Ben Jones · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“Glove: Global vectors for word representation”
Jeffrey Pennington, Richard Socher and Christopher Manning · 2014
Earlier work this paper cites.
“Academic libraries and discovery tools: A survey of the literature”
Beth Thomsett-Scott and Patricia Reese · 2016
Earlier work this paper cites.
“The Changing Landscape of Research Library Collections: Ensuring Realistic Sustainability.”, 2016
Daniel Tsang and Julia Gelfand · 2016
Earlier work this paper cites.
“Towards A Rigorous Science of Interpretable Machine Learning”, 2017
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Cited alongside, same era.
“The Unified Astronomy Thesaurus: Semantic metadata for astronomy and astrophysics”
Katie Frey and Alberto Accomazzi · 2018
Cited alongside, same era.
“Deep contextualized word representations”
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee and Luke Zettlemoyer · 2018
Cited alongside, same era.
“Learning and Evaluating Sparse Interpretable Sentence Embeddings”, 2018
Valentin Trifonov, Octavian-Eugen Ganea, Anna Potapenko and Thomas Hofmann · 2018
Cited alongside, same era.
“Linguistic knowledge and transferability of contextual representations”
“Step-gs: Guiding large language models via step-by-step prompting”
Yichen Cao, Xinyi Wang, Yiran Cao, Renfeng Xu, Zhihan Dong, Qi Fang, Yeyun Gong, Lei Li, Shuming Shi and Jiafeng Yan · 2023
Later among the works it cites.
“Sparse Autoencoders Find Highly Interpretable Features in Language Models”, 2023
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben and Lee Sharkey · 2023
Later among the works it cites.
“Sparse autoencoders find highly interpretable features in language models”
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben and Lee Sharkey · 2023
Later among the works it cites.
“Neuron to graph: Interpreting language model neurons at scale”
Alex Foote, Neel Nanda, Esben Kran, Ioannis Konstas, Shay Cohen and Fazl Barez · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nelson Liu, Matt Gardner, Yonatan Belinkov, Matthew Peters and Noah Smith · 2019
Cited alongside, same era.
“Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks”
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
“Evaluating Gender Bias in Machine Translation”
Gabriel Stanovsky, Noah Smith and Luke Zettlemoyer · 2019
Cited alongside, same era.
“An analysis of cross-domain performance in deep neural networks”
Nicholas Thompson, Kristjan Greenewald, Keegan Lee and Gabriel Manso · 2020
Cited alongside, same era.
“The evolving role of library collections in the broader information ecosystem”
Mark Dahl · 2021
Cited alongside, same era.
“SimCSE: Simple contrastive learning of sentence embeddings”
Tianyu Gao, Xingcheng Yao and Danqi Chen · 2021
Cited alongside, same era.
“Toy Models of Superposition”, 2022
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg and Christopher Olah · 2022
Cited alongside, same era.
“Softmax Linear Units”, 2022
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Johnston, Ben Mann, Amanda Askell, Danny Hernandez, Dawn Drain and Zac Hatfield-Dodds · 2022
Cited alongside, same era.
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii and Dimitris Bertsimas · 2023
Later among the works it cites.
“Open Source Replication & Commentary on Anthropic’s Dictionary Learning Paper” [Accessed 22-07-2024], 2023
Neel Nanda · 2023
Later among the works it cites.
“Activation Steering with SAEs” Accessed 16-07-2024, 2024
Arthur Conmy and Neel Nanda · 2024
Closest in time.
“Interpreting and Steering Features in Images” [Accessed 16-07-2024],
Gytis Daujotas · 2024
Closest in time.
“Not All Language Model Features Are Linear”, 2024
Joshua Engels, Isaac Liao, Eric. Michaud, Wes Gurnee and Max Tegmark · 2024
Closest in time.
“Scaling Laws for Neurons in GPT Models”
Leo Gao, John Thickstun, Anirudh Madaan, Zach Scherlis, Arush Guha, Sumanth Dathathri, Jared Kaplan, Azalia Mirhoseini and Ilya Sutskever · 2024
Closest in time.
“Ghost Grads: An improvement on resampling” [Accessed 19-07-2024], 2023
Adam Jermyn and Adly Templeton · 2024
Closest in time.
“Prism: mapping interpretable concepts and features in a latent space of language” Accessed 16-07-2024, 2024
Linus Lee · 2024
Closest in time.
“Towards principled evaluations of sparse autoencoders for interpretability and control”
Aleksandar Makelov, George Lange and Neel Nanda · 2024
Closest in time.
“Improving dictionary learning with gated sparse autoencoders”
Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah and Neel Nanda · 2024
Closest in time.
Zechang Sun, Yuan-Sen Ting, Yaobo Liang, Nan Duan, Song Huang and Zheng Cai · 2024
Closest in time.
“Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet”
Adly Templeton · 2024
Closest in time.
“Text Embeddings by Weakly-Supervised Contrastive Pre-training”, 2024
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder and Furu Wei · 2024
Closest in time.
“Addressing Feature Suppression in SAEs” [Accessed 16-07-2024],
Benjamin Wright and Lee Sharkey · 2024
Closest in time.