Fetching the paper…
Reading the bibliography…
The Universality Hypothesis in large language models (LLMs) claims that different models converge towards similar concept representations in their latent spaces.
Similarity of neural network representations revisited, 2019
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton · 1905
Earlier work this paper cites.
Relations between two sets of variates
Harold Hotelling · 1936
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A. Olshausen and David J. Field · 1997
Earlier work this paper cites.
Representational similarity analysis – connecting the branches of systems neuroscience
Nikolaus Kriegeskorte, Marieke Mur, and Peter Bandettini · 2008
Earlier work this paper cites.
Alireza Makhzani and Brendan Frey · 2013
Earlier work this paper cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson · 2014
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Narain Sohl-Dickstein · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Towards understanding learning representations: To what extent do different neural networks learn the same representation
Liwei Wang, Lunjia Hu, Jiayuan Gu, Yue Wu, Zhiqiang Hu, Kun He, and John Hopcroft · 2018
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
High-low frequency detectors
Ludwig Schubert, Chelsea Voss, Nick Cammarata, Gabriel Goh, and Chris Olah · 2020
Earlier work this paper cites.
Revisiting model stitching to compare neural representations
Yamini Bansal, Preetum Nakkiran, and Boaz Barak · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Earlier work this paper cites.
The Alignment Problem from a Deep Learning Perspective, 2022
Richard Ngo, Lawrence Chan, and Sören Mindermann · 2022
Earlier work this paper cites.
Goal Misgeneralization: Why Correct Specifications Aren’t Enough For Correct Goals, November 2022
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton · 2022
Earlier work this paper cites.
Interim research report: Taking features out of superposition with sparse autoencoders
Lee Sharkey, Dan Braun, and Beren Millidge · 2022
Earlier work this paper cites.
System iii: Learning with domain knowledge for safety constraints, 2023
Fazl Barez, Hosien Hasanbieg, and Alesandro Abbate · 2023
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling, 2023
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Earlier work this paper cites.
A toy model of universality: reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Earlier work this paper cites.
Sparse autoencoders find highly interpretable features in language models, 2023
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Earlier work this paper cites.
Bipartite invariance in mouse primary visual cortex
Zhiwei Ding, Dat T. Tran, Kayla Ponder, Erick Cobos, Zhuokun Ding, Paul G. Fahey, Eric Wang, Taliah Muhammad, Jiakun Fu, Santiago A. Cadena, Stelios Papadopoulos, Saumil Patel, Katrin Franke, Jacob Reimer, Fabian H. Sinz, Alexander S. Ecker, Xaq Pitkow, and Andreas S. Tolias · 2023
Earlier work this paper cites.
Tinystories-1layer-21m
Ronen Eldan · 2023
Earlier work this paper cites.
Sparse autoencoders (sae) repository
EleutherAI · 2023
Cited alongside, same era.
Neuron to graph: Interpreting language model neurons at scale, 2023
Alex Foote, Neel Nanda, Esben Kran, Ioannis Konstas, Shay Cohen, and Fazl Barez · 2023
Cited alongside, same era.
Deepdecipher: Accessing and investigating neuron activation in large language models, 2023
Albert Garde, Esben Kran, and Fazl Barez · 2023
Cited alongside, same era.
An overview of catastrophic ai risks, 2023
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside · 2023
Cited alongside, same era.
Similarity of neural network models: A survey of functional and representational measures, 2023
Max Klabunde, Tobias Schumacher, Markus Strohmaier, and Florian Lemmerich · 2023
Steering language model refusal with sparse autoencoders, 2024
Kyle O’Brien, David Majercak, Xavier Fernandes, Richard Edgar, Jingya Chen, Harsha Nori, Dean Carignan, Eric Horvitz, and Forough Poursabzi-Sangde · 2024
Closest in time.
Steering llama 2 via contrastive activation addition, 2024
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner · 2024
Closest in time.
A list of 45 mech interp project ideas from apollo research
Lee Sharkey, Lucius Bushnaq, Dan Braun, Stefan Hex, and Nicholas Goldowsky-Dill · 2024
Closest in time.
Large concept models: Language modeling in a sentence representation space, 2024
LCM team, Loïc Barrault, Paul-Ambroise Duquenne, Maha Elbayad, Artyom Kozhevnikov, Belen Alastruey, Pierre Andrews, Mariano Coria, Guillaume Couairon, Marta R. Costa-jussà, David Dale, Hady Elsahar, Kevin Heffernan, João Maria Janeiro, Tuan Tran, Christophe Ropers, Eduardo Sánchez, Robin San Roman, Alexandre Mourachko, Safiyyah Saleem, and Holger Schwenk · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Linearly mapping from image to text space, 2023
Jack Merullo, Louis Castricato, Carsten Eickhoff, and Ellie Pavlick · 2023
Cited alongside, same era.
Understanding and Controlling a Maze-Solving Policy Network, 2023
Ulisse Mini, Peli Grietzer, Mrinank Sharma, Austin Meek, Monte MacDiarmid, and Alexander Matt Turner · 2023
Cited alongside, same era.
Feature manifold toy model
Chris Olah and Josh Batson · 2023
Cited alongside, same era.
Getting aligned on representational alignment
Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller, Andi Peng, Andreea Bobu, Been Kim, Bradley C. Love, Erin Grant, Jascha Achterberg, Joshua B. Tenenbaum, Katherine M. Collins, Katherine L. Hermann, Kerem Oktar, Klaus Greff, Martin N. Hebart, Nori Jacoby, Qiuyi Zhang, Raja Marjieh, Robert Geirhos, Sherol Chen, Simon Kornblith, Sunayana Rane, Talia Konkle, Thomas P. O’Connell, Thomas Unterthiner, Andrew K. Lampinen, Klaus-Robert Müller, Mariya Toneva, and Thomas L. Griffiths · 2023
Cited alongside, same era.
Representation engineering: A top-down approach to ai transparency, 2023
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks · 2023
Cited alongside, same era.
Refusal in language models is mediated by a single direction, 2024
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda · 2024
Cited alongside, same era.
Mechanistic interpretability for ai safety – a review, 2024
Leonard Bereska and Efstratios Gavves · 2024
Cited alongside, same era.
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.
Steering language models with activation engineering, 2024
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid · 2024
Closest in time.
Redpajama: an open dataset for training large language models, 2024
Maurice Weber, Daniel Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Ré, Irina Rish, and Ce Zhang · 2024
Closest in time.
Sparse autoencoders for dense text embeddings reveal hierarchical feature sub-structure
Christine Ye, Charles O’Neill, John F Wu, and Kartheik G. Iyer · 2024
Closest in time.
Not all language model features are one-dimensionally linear, 2025
Joshua Engels, Eric J. Michaud, Isaac Liao, Wes Gurnee, and Max Tegmark · 2025
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupre la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2025
Closest in time.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2025
Closest in time.
Sparse autoencoders can interpret randomly initialized transformers, 2025
Thomas Heap, Tim Lawson, Lucy Farnik, and Laurence Aitchison · 2025
Closest in time.
Projecting assumptions: The duality between sparse autoencoders and concept geometry, 2025
Sai Sumedh R. Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba Ba · 2025
Closest in time.
Saebench: A comprehensive benchmark for sparse autoencoders, 2024
A. Karvonen, C. Rager, J. Lin, C. Tigges, J. Bloom, D. Chanin, Y.-T. Lau, E. Farrell, A. Conmy, C. McDougall, K. Ayonrinde, M. Wearden, S. Marks, and N. Nanda · 2025
Closest in time.
SAEs (usually) Transfer Between Base and Chat Models, 2024a
Connor Kissane, Robert Krzyzanowski, Arthur Conmy, and Neel Nanda · 2025
Closest in time.
Saes are highly dataset dependent: a case study on the refusal direction, November 2024b
Connor Kissane, robertzk, Neel Nanda, and Arthur Conmy · 2025
Closest in time.
Representational similarity via interpretable visual concepts
Neehar Kondapaneni, Oisin Mac Aodha, and Pietro Perona · 2025
Closest in time.
Do sparse autoencoders (saes) transfer across base and finetuned language models?, sep 2024
Taras Kutsyk, Tommaso Mencattini, and Ciprian Florea · 2025
Closest in time.
Calendar Feature Geometry in GPT-2 Layer 8 Residual Stream SAEs
Patrick Leask, Bart Bussmann, and Neel Nanda · 2025
Closest in time.
Sparse autoencoders do not find canonical units of analysis
Patrick Leask, Bart Bussmann, Michael T. Pearce, Joseph I. Bloom, Curt Tigges, N. Al Moubayed, Lee Sharkey, and Neel Nanda · 2025
Closest in time.
The geometry of concepts: Sparse autoencoder feature structure
Yuxiao Li, Eric J. Michaud, David D. Baek, Joshua Engels, Xiaoqing Sun, and Max Tegmark · 2025
Closest in time.
Insights on crosscoder model diffing
Siddharth Mishra-Sharma, Trenton Bricken, Jack Lindsey, Adam Jermyn, Jonathan Marcus, Kelley Rivoire, Christopher Olah, and Thomas Henighan · 2025
Closest in time.
Sparse autoencoders trained on the same data learn different features, 2025
Gonçalo Paulo and Nora Belrose · 2025
Closest in time.
Universal sparse autoencoders: Interpretable cross-model concept alignment, 2025
Harrish Thasarathan, Julian Forsyth, Thomas Fel, Matthew Kowal, and Konstantinos Derpanis · 2025
Closest in time.
Do sparse autoencoders find ’true features’?, 2023
Demian Till · 2025
Closest in time.