Fetching the paper…
Reading the bibliography…
We analyze identifiability as a possible explanation for the ubiquity of linear properties across language models, such as the vector difference between the representations of "easy" and "easiest" being parallel to that between "lucky" and "luckiest".
Fundamental ideas and methods of the theory of relativity, presented in their development
Albert Einstein · 1920
Earlier work this paper cites.
A model for analogical reasoning
David E. Rumelhart and Adele A. Abrahamson · 1973
Earlier work this paper cites.
When do projections commute?
W Rehder · 1980
Earlier work this paper cites.
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2000
Earlier work this paper cites.
Gravitation
Kip S. Thorne, Charles W. Misner, and John Archibald Wheeler · 2000
Earlier work this paper cites.
Learning distributed representations of concepts using linear relational embedding
Alberto Paccanaro and Geoffrey E. Hinton · 2001
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Linear Algebra Done Right
Sheldon Axler · 2015
Earlier work this paper cites.
A latent variable model approach to PMI-based word embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2016
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain · 2016
Earlier work this paper cites.
Modern Classical Physics: Optics, Fluids, Plasmas, Elasticity, Relativity, and Statistical Physics
Kip S. Thorne and Roger D. Blandford · 2017
Earlier work this paper cites.
Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
Generating wikipedia by summarizing long sequences
Peter J Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer · 2018
Earlier work this paper cites.
Analogies explained: Towards understanding word embeddings
Carl Allen and Timothy Hospedales · 2019
Earlier work this paper cites.
Nonlinear ICA using auxiliary variables and generalized contrastive learning
Aapo Hyvarinen, Hiroaki Sasaki, and Richard Turner · 2019
Earlier work this paper cites.
The incomplete Rosetta Stone problem: Identifiability results for multi-view nonlinear ICA
Luigi Gresele, Paul K Rubenstein, Arash Mehrjou, Francesco Locatello, and Bernhard Schölkopf · 2020
Earlier work this paper cites.
Hidden markov nonlinear ICA: Unsupervised learning from nonstationary time series
Hermanni Hälvä and Aapo Hyvarinen · 2020
Earlier work this paper cites.
On linear identifiability of learned representations
Geoffrey Roeder, Luke Metz, and Durk Kingma · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki · 2021
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Cited alongside, same era.
The geometry of multilingual language model representations
Tyler A. Chang, Zhuowen Tu, and Benjamin K Bergen · 2022
Cited alongside, same era.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
Indeterminacy in generative models: Characterization and strong identifiability
Quanhan Xi and Benjamin Bloem-Reddy · 2023
Later among the works it cites.
Interventional causal representation learning
Kartik Ahuja, Divyat Mahajan, Yixin Wang, and Yoshua Bengio · 2023
Later among the works it cites.
Linearity of relation decoding in transformer language models
Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau · 2024
Closest in time.
Monotonic representation of numeric properties in language models
Benjamin Heinzerling and Kentaro Inui · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2022
Cited alongside, same era.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Cited alongside, same era.
Identifiable deep generative models via sparse decoding
Gemma E. Moran, Dhanya Sridhar, Yixin Wang, and David M. Blei · 2022
Cited alongside, same era.
Function classes for identifiable nonlinear independent component analysis
Simon Buchholz, Michel Besserve, and Bernhard Schölkopf · 2022
Cited alongside, same era.
Binary independent component analysis: a non-stationarity-based approach
Antti Hyttinen, Vitória Barin Pacela, and Aapo Hyvärinen · 2022
Cited alongside, same era.
Citris: Causal identifiability from temporal intervened sequences
Phillip Lippe, Sara Magliacane, Sindy Löwe, Yuki M Asano, Taco Cohen, and Stratis Gavves · 2022
Cited alongside, same era.
Language models implement simple word2vec-style vector arithmetic
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2023
Cited alongside, same era.
Challenges in explaining representational similarity through identifiability
Beatrix Miranda Ginn Nielsen, Luigi Gresele, and Andrea Dittadi · 2024
Closest in time.
On the origins of linear representations in large language models
Yibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon Aragam, and Victor Veitch · 2024
Closest in time.
Improving instruction-following in language models through activation steering
Alessandro Stolfo, Vidhisha Balachandran, Safoora Yousefi, Eric Horvitz, and Besmira Nushi · 2024
Closest in time.
Not all language model features are linear
Joshua Engels, Isaac Liao, Eric J Michaud, Wes Gurnee, and Max Tegmark · 2024
Closest in time.
Causal Component Analysis
Wendong Liang, Armin Kekić, Julius von Kügelgen, Simon Buchholz, Michel Besserve, Luigi Gresele, and Bernhard Schölkopf · 2024
Closest in time.
Nonparametric identifiability of causal representations from unknown interventions
Julius von Kügelgen, Michel Besserve, Liang Wendong, Luigi Gresele, Armin Kekić, Elias Bareinboim, David Blei, and Bernhard Schölkopf · 2024
Closest in time.
General identifiability and achievability for causal representation learning
Burak Varici, Emre Acartürk, Karthikeyan Shanmugam, and Ali Tajer · 2024
Closest in time.
Learning interpretable concepts: Unifying causal representation learning and foundation models
Goutham Rajendran, Simon Buchholz, Bryon Aragam, Bernhard Schölkopf, and Pradeep Ravikumar · 2024
Closest in time.
Disentangled representation learning in non-markovian causal systems
Adam Li, Yushu Pan, and Elias Bareinboim · 2024
Closest in time.
Learning partitions from context
Simon Buchholz · 2024
Closest in time.
Position: Understanding LLMs Requires More Than Statistical Generalization
Patrik Reizinger, Szilvia Ujváry, Anna Mészáros, Anna Kerekes, Wieland Brendel, and Ferenc Huszár · 2024
Closest in time.
Does localization inform editing? Surprising differences in causality-based localization vs. knowledge editing in language models
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun · 2024
Closest in time.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda · 2024
Closest in time.
Shortcuts and identifiability in concept-based models from a neuro-symbolic lens
Samuele Bortolotti, Emanuele Marconato, Paolo Morettin, Andrea Passerini, and Stefano Teso · 2025
Closest in time.