Fetching the paper…
Reading the bibliography…
Autoencoders have been used for finding interpretable and disentangled features underlying neural network representations in both image and text domains.
How to not measure disentanglement
Anna Sepliarskaia, Julia Kiseleva, and Maarten de Rijke. 2019 · 1910
Earlier work this paper cites.
Weakly supervised disentanglement with guarantees
Rui Shu, Yining Chen, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2019 · 1910
Earlier work this paper cites.
Introduction to the theory of computation
Michael Sipser. 1996 · 1996
Earlier work this paper cites.
Nonlinear independent component analysis: Existence and uniqueness results
Aapo Hyvärinen and Petteri Pajunen. 1999 · 1999
Earlier work this paper cites.
Towards nonlinear disentanglement in natural data with temporal sparse coding
David Klindt, Lukas Schott, Yash Sharma, Ivan Ustyuzhaninov, Wieland Brendel, Matthias Bethge, and Dylan Paiton. 2020 · 2007
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal. 2020 · 2009
Earlier work this paper cites.
Causalworld: A robotic manipulation benchmark for causal structure and transfer learning
Ossama Ahmed, Frederik Träuble, Anirudh Goyal, Alexander Neitz, Yoshua Bengio, Bernhard Schölkopf, Manuel Wüthrich, and Stefan Bauer. 2020 · 2010
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. 2017 · 2017
Earlier work this paper cites.
Nonlinear ICA of temporally dependent stationary sources
Aapo Hyvarinen and Hiroshi Morioka. 2017 · 2017
Earlier work this paper cites.
Nonlinear ICA using auxiliary variables and generalized contrastive learning
Aapo Hyvarinen, Hiroaki Sasaki, and Richard Turner. 2019 · 2019
Earlier work this paper cites.
Challenging common assumptions in the unsupervised learning of disentangled representations
Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem. 2019 · 2019
Earlier work this paper cites.
Manifold mixup: Better representations by interpolating hidden states
Vikas Verma, Alex Lamb, Christopher Beckham, Amir Najafi, Ioannis Mitliagkas, David Lopez-Paz, and Yoshua Bengio. 2019 · 2019
Earlier work this paper cites.
Weakly-supervised disentanglement without compromises
Francesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf, Olivier Bachem, and Michael Tschannen. 2020 · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021 · 2021
Earlier work this paper cites.
On disentangled representations learned from correlated data
Frederik Träuble, Elliot Creager, Niki Kilbertus, Francesco Locatello, Andrea Dittadi, Anirudh Goyal, Bernhard Schölkopf, and Stefan Bauer. 2021 · 2021
Cited alongside, same era.
Weakly supervised causal representation learning
Johann Brehmer, Pim De Haan, Phillip Lippe, and Taco Cohen. 2022 · 2022
Cited alongside, same era.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2022 · 2022
Cited alongside, same era.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. 2022 · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
How deep neural networks learn compositional data: The random hierarchy model
Francesco Cagnetta, Leonardo Petrini, Umberto M Tomasini, Alessandro Favero, and Matthieu Wyart. 2024 · 2024
Closest in time.
Sparse autoencoders reveal temporal difference learning in large language models
Can Demircan, Tankred Saanum, Akshay K. Jagadish, Marcel Binz, and Eric Schulz. 2024 · 2024
Closest in time.
Not all language model features are linear
Joshua Engels, Isaac Liao, Eric J Michaud, Wes Gurnee, and Max Tegmark. 2024 · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. 2024 · 2024
Closest in time.
Sparse autoencoders reveal universal feature spaces across large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2022 · 2022
Cited alongside, same era.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. 2023 · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L. Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E. Burke, Tristan Hume, Shan Carter, Tom Henighan, and Chris Olah. 2023 · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. 2023 · 2023
Cited alongside, same era.
Synergies between disentanglement and sparsity: Generalization and identifiability in multi-task learning
Sébastien Lachapelle, Tristan Deleu, Divyat Mahajan, Ioannis Mitliagkas, Yoshua Bengio, Simon Lacoste-Julien, and Quentin Bertrand. 2023 · 2023
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. 2023 · 2023
Cited alongside, same era.
Linear causal disentanglement via interventions
Chandler Squires, Anna Seigal, Salil S Bhate, and Caroline Uhler. 2023 · 2023
Cited alongside, same era.
Physics of language models: Part 1, learning hierarchical language structures
Zeyuan Allen-Zhu and Yuanzhi Li. 2024 · 2024
Cited alongside, same era.
Michael Lan, Philip Torr, Austin Meek, Ashkan Khakzar, David Krueger, and Fazl Barez. 2024 · 2024
Closest in time.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2
Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca Dragan, Rohin Shah, and Neel Nanda. 2024 · 2024
Closest in time.
A percolation model of emergence: Analyzing transformers trained on a formal language
Ekdeep Singh Lubana, Kyogo Kawaguchi, Robert P. Dick, and Hidenori Tanaka. 2024 · 2024
Closest in time.
Towards principled evaluations of sparse autoencoders for interpretability and control
Aleksandar Makelov, George Lange, and Neel Nanda. 2024 · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller. 2024 · 2024
Closest in time.
Improving dictionary learning with gated sparse autoencoders
Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah, and Neel Nanda. 2024 · 2024
Closest in time.
[interim research report] taking features out of superposition with sparse autoencoders
Lee Sharkey, Dan Braun, and Beren. 2024 · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. 2024 · 2024
Closest in time.
Transformers are uninterpretable with myopic methods: a case study with bounded dyck grammars
Kaiyue Wen, Yuchen Li, Bingbin Liu, and Andrej Risteski. 2024 · 2024
Closest in time.
Axbench: Steering llms? even simple baselines outperform sparse autoencoders
Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher D Manning, and Christopher Potts. 2025 · 2025
Closest in time.