Fetching the paper…
Reading the bibliography…
Recent work has proposed that language models perform computation by manipulating one-dimensional representations of concepts ("features") in activation space.
über die abgrenzung der eigenwerte einer matrix
Semyon Aranovich Gershgorin · 1931
Earlier work this paper cites.
Extensions of lipschitz mappings into a hilbert space
William B. Johnson and Joram Lindenstrauss · 1984
Earlier work this paper cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 1984
Earlier work this paper cites.
Problems and results in extremal combinatorics—i
Noga Alon · 2003
Earlier work this paper cites.
Almost orthogonal vectors
Bill Johnson (https://mathoverflow.net/users/2554/bill johnson) · 2010
Earlier work this paper cites.
A Mathematical Introduction to Compressive Sensing
Simon Foucart and Holger Rauhut · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomáš Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
A framework for parallelizing hierarchical clustering methods
Silvio Lattanzi, Thomas Lavastida, Kefu Lu, and Benjamin Moseley · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Singular value inequalities
Nicholas J. Higham · 2021
Earlier work this paper cites.
Interpreting neural networks through the polytope lens
Sid Black, Lee Sharkey, Leo Grinsztajn, Eric Winsor, Dan Braun, Jacob Merizian, Kip Parker, Carlos Ramón Guevara, Beren Millidge, Gabriel Alfour, et al · 2022
Earlier work this paper cites.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Earlier work this paper cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2022
Earlier work this paper cites.
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas S Nolte, Eric Michaud, Max Tegmark, and Mike Williams · 2022
Earlier work this paper cites.
Transformerlens
Neel Nanda and Joseph Bloom · 2022
Earlier work this paper cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Cited alongside, same era.
Towards best practices of activation patching in language models: Metrics and methods
Fred Zhang and Neel Nanda · 2023
Later among the works it cites.
Llama 3 model card, 2024
AI@Meta · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Closest in time.
Open source sparse autoencoders for all residual stream layers of gpt2 small
Joseph Bloom · 2024
Closest in time.
Recurrent neural networks learn to store and generate sequences using non-linear representations
Róbert Csordás, Christopher Potts, Christopher D Manning, and Atticus Geiger · 2024
Closest in time.
Towards guaranteed safe ai: A framework for ensuring robust and reliable ai systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Cited alongside, same era.
Successor heads: Recurring, interpretable attention heads in the wild
Rhys Gould, Euan Ong, George Ogden, and Arthur Conmy · 2023
Cited alongside, same era.
Language models represent space and time
Wes Gurnee and Max Tegmark · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Samuel Marks and Max Tegmark · 2023
Cited alongside, same era.
Feature emergence via margin maximization: case studies in algebraic tasks
Depen Morwani, Benjamin L Edelman, Costin-Andrei Oncescu, Rosie Zhao, and Sham Kakade · 2023
Cited alongside, same era.
David Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, et al · 2024
Closest in time.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Michael Hanna, Ollie Liu, and Alexandre Variengien · 2024
Closest in time.
Monotonic representation of numeric properties in language models
Benjamin Heinzerling and Kentaro Inui · 2024
Closest in time.
On the origins of linear representations in large language models
Yibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon Aragam, and Victor Veitch · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Sae feature geometry is outside the superposition hypothesis
Jake Mendel · 2024
Closest in time.
Opening the ai black box: program synthesis via mechanistic interpretability
Eric J Michaud, Isaac Liao, Vedang Lad, Ziming Liu, Anish Mudide, Chloe Loughridge, Zifan Carl Guo, Tara Rezaei Kheirkhah, Mateja Vukelić, and Max Tegmark · 2024
Closest in time.
Aaron Mueller, Jannik Brinkmann, Millicent Li, Samuel Marks, Koyena Pal, Nikhil Prakash, Can Rager, Aruna Sankaranarayanan, Arnab Sen Sharma, Jiuding Sun, et al · 2024
Closest in time.
What is a linear representation? what is a multidimensional feature?
Chris Olah · 2024
Closest in time.
Transformers represent belief state geometry in their residual stream
Adam Shai, Paul Riechers, Lucas Teixeira, Alexander Oldenziel, and Sarah Marzen · 2024
Closest in time.
Gpt-2’s positional embedding matrix is a helix, 2023a
Adam Yedidia · 2024
Closest in time.
The positional embedding matrix and previous-token heads: how do they actually work?, 2023b
Adam Yedidia · 2024
Closest in time.
Language models use trigonometry to do addition
Subhash Kantamneni and Max Tegmark · 2025
Closest in time.