Fetching the paper…
Reading the bibliography…
The linear representation hypothesis is the informal idea that semantic concepts are encoded as linear directions in the representation spaces of large language models (LLMs).
Wordnet: a lexical database for english
George A Miller · 1995
Earlier work this paper cites.
A well-conditioned estimator for large-dimensional covariance matrices
Olivier Ledoit and Michael Wolf · 2004
Earlier work this paper cites.
Dynamic topic models
David M Blei and John D Lafferty · 2006
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomáš Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
A latent variable model approach to pmi-based word embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2016
Earlier work this paper cites.
Exponential family embeddings
Maja Rudolph, Francisco Ruiz, Stephan Mandt, and David Blei · 2016
Earlier work this paper cites.
Skip-gram - zipf + uniform = vector additivity
Alex Gittens, Dimitris Achlioptas, and Michael W. Mahoney · 2017
Earlier work this paper cites.
The strange geometry of skip-gram with negative sampling
David Mimno and Laure Thompson · 2017
Earlier work this paper cites.
Poincaré embeddings for learning hierarchical representations
Maximillian Nickel and Douwe Kiela · 2017
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2018
Earlier work this paper cites.
Towards understanding linear word analogies
Kawin Ethayarajh, David Duvenaud, and Graeme Hirst · 2018
Earlier work this paper cites.
Hyperbolic entailment cones for learning hierarchical embeddings
Octavian Ganea, Gary Bécigneul, and Thomas Hofmann · 2018
Earlier work this paper cites.
Analogies explained: Towards understanding word embeddings
Carl Allen and Timothy Hospedales · 2019
Earlier work this paper cites.
Understanding composition of word embeddings via tensor decomposition
Abraham Frandsen and Rong Ge · 2019
Earlier work this paper cites.
Visualizing and measuring the geometry of bert
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim · 2019
Cited alongside, same era.
On the sentence embeddings from pre-trained language models
Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li · 2020
Cited alongside, same era.
Probing bert in hyperbolic spaces
Boli Chen, Yao Fu, Guangwei Xu, Pengjun Xie, Chuanqi Tan, Mosha Chen, and Liping Jing · 2021
Cited alongside, same era.
Natural alpha embeddings
Riccardo Volpi and Luigi Malagò · 2021
Cited alongside, same era.
A non-graphical representation of conditional independence via the neighbourhood lattice
Arash A Amini, Bryon Aragam, and Qing Zhou · 2022
Cited alongside, same era.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Linear representations of sentiment in large language models
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda · 2023
Later among the works it cites.
Concept algebra for (score-based) text-controlled generative models
Zihao Wang, Lin Gui, Jeffrey Negrea, and Victor Veitch · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Cited alongside, same era.
The geometry of multilingual language model representations
Tyler A Chang, Zhuowen Tu, and Benjamin K Bergen · 2022
Cited alongside, same era.
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, et al · 2022
Cited alongside, same era.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Victor Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Y Zou · 2022
Cited alongside, same era.
Relative representations enable zero-shot latent space communication
Luca Moschella, Valentino Maiorca, Marco Fumero, Antonio Norelli, Francesco Locatello, and Emanuele Rodola · 2022
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Cited alongside, same era.
Later among the works it cites.
Identifying functionally important features with end-to-end sparse dictionary learning
Dan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, and Lee Sharkey · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Language models represent space and time
Wes Gurnee and Max Tegmark · 2024
Closest in time.
Language models as hierarchy encoders
Yuan He, Zhangdie Yuan, Jiaoyan Chen, and Ian Horrocks · 2024
Closest in time.
On the origins of linear representations in large language models
Yibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon Aragam, and Victor Veitch · 2024
Closest in time.
Sparse autoencoders work on attention layer outputs
Connor Kissane, Robert Krzyzanowski, Arthur Conmy, and Neel Nanda · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2024
Closest in time.
The geometry of hidden representations of large transformer models
Lucrezia Valeriani, Diego Doimo, Francesca Cuturello, Alessandro Laio, Alessio Ansuini, and Alberto Cazzaniga · 2024
Closest in time.