Fetching the paper…
Reading the bibliography…
Recent works have argued that high-level semantic concepts are encoded "linearly" in the representation space of large language models.
“Learning mixtures of Gaussians”
Sanjoy Dasgupta · 1999
Earlier work this paper cites.
“Causation, prediction, and search”
Peter Spirtes, Clark Glymour and Richard Scheines · 2000
Earlier work this paper cites.
“Stochastic approximation and recursive algorithms and applications”
Harold Kushner and G Yin · 2003
Earlier work this paper cites.
“Dynamic topic models”
David Blei and John Lafferty · 2006
Earlier work this paper cites.
“Probabilistic graphical models: principles and techniques”
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
“Causality”
Judea Pearl · 2009
Earlier work this paper cites.
“The winograd schema challenge”
Hector Levesque, Ernest Davis and Leora Morgenstern · 2012
Earlier work this paper cites.
“Parallel Data, Tools and Interfaces in OPUS”
Jörg Tiedemann · 2012
Earlier work this paper cites.
“A Practical Algorithm for Topic Modeling with Provable Guarantees”
Sanjeev Arora et al · 2013
Earlier work this paper cites.
“Linguistic regularities in continuous space word representations”
Tom\’as Mikolov, Wen-tau Yih and Geoffrey Zweig · 2013
Earlier work this paper cites.
“Glove: Global vectors for word representation”
Jeffrey Pennington, Richard Socher and Christopher Manning · 2014
Earlier work this paper cites.
Sanjeev Arora et al · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Unsupervised representation learning with deep convolutional generative adversarial networks”
Alec Radford, Luke Metz and Soumith Chintala · 2015
Earlier work this paper cites.
“A latent variable model approach to pmi-based word embeddings”
Sanjeev Arora et al · 2016
Earlier work this paper cites.
“Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn’t.”
Anna Gladkova, Aleksandr Drozd and Satoshi Matsuoka · 2016
Earlier work this paper cites.
“Exponential family embeddings”
Maja Rudolph, Francisco Ruiz, Stephan Mandt and David Blei · 2016
Earlier work this paper cites.
“Network dissection: Quantifying interpretability of deep visual representations”
David Bau et al · 2017
Earlier work this paper cites.
“Latent constraints: Learning to generate conditionally from unconditional generative models”
Jesse Engel, Matthew Hoffman and Adam Roberts · 2017
Earlier work this paper cites.
“Skip-Gram - Zipf + Uniform = Vector Additivity”
Alex Gittens, Dimitris Achlioptas and Michael. Mahoney · 2017
Earlier work this paper cites.
“Skip-gram- zipf+ uniform= vector additivity”
Alex Gittens, Dimitris Achlioptas and Michael Mahoney · 2017
Earlier work this paper cites.
“The strange geometry of skip-gram with negative sampling”
David Mimno and Laure Thompson · 2017
Earlier work this paper cites.
“Svcca: Singular vector canonical correlation analysis for deep understanding and improvement”
Maithra Raghu, Justin Gilmer, Jason Yosinski and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
“Dynamic bernoulli embeddings for language evolution”
Maja Rudolph and David Blei · 2017
Earlier work this paper cites.
“Linear algebraic structure of word senses, with applications to polysemy”
Sanjeev Arora et al · 2018
Cited alongside, same era.
“Towards understanding linear word analogies”
Kawin Ethayarajh, David Duvenaud and Graeme Hirst · 2018
Cited alongside, same era.
“Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)”
Been Kim et al · 2018
Cited alongside, same era.
“The implicit bias of gradient descent on separable data”
Daniel Soudry et al · 2018
Cited alongside, same era.
“What the vec? towards probabilistically grounded embeddings”
Carl Allen, Ivana Balazevic and Timothy Hospedales · 2019
Cited alongside, same era.
“Analogies explained: Towards understanding word embeddings”
“Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning”
Victor Liang et al · 2022
Later among the works it cites.
“Acquisition of chess knowledge in alphazero”
Thomas McGrath et al · 2022
Later among the works it cites.
“Relative representations enable zero-shot latent space communication”
Luca Moschella et al · 2022
Later among the works it cites.
“From statistical to causal learning”
Bernhard Sch\"olkopf and Julius von K\"ugelgen · 2022
Later among the works it cites.
“Causal structure learning: a combinatorial perspective”
Chandler Squires and Caroline Uhler · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carl Allen and Timothy Hospedales · 2019
Cited alongside, same era.
“Understanding composition of word embeddings via tensor decomposition”
Abraham Frandsen and Rong Ge · 2019
Cited alongside, same era.
“Visualizing and measuring the geometry of BERT”
Emily Reif et al · 2019
Cited alongside, same era.
“Variational autoencoders and nonlinear ica: A unifying framework”
Ilyes Khemakhem, Diederik Kingma, Ricardo Monti and Aapo Hyvarinen · 2020
Cited alongside, same era.
“On the sentence embeddings from pre-trained language models”
Bohan Li et al · 2020
Cited alongside, same era.
“Evaluating Natural Alpha Embeddings on Intrinsic and Extrinsic Tasks”
Riccardo Volpi and Luigi Malag\‘o · 2020
Cited alongside, same era.
“Direction matters: On the implicit bias of stochastic gradient descent with moderate learning rate”
Jingfeng Wu, Difan Zou, Vladimir Braverman and Quanquan Gu · 2020
Cited alongside, same era.
Simon Buchholz et al · 2023
Later among the works it cites.
“Finding Neurons in a Haystack: Case Studies with Sparse Probing”
Wes Gurnee et al · 2023
Later among the works it cites.
“Identifiability of latent-variable and structural-equation models: from linear to nonlinear”
Aapo Hyv\"arinen, Ilyes Khemakhem and Ricardo Monti · 2023
Later among the works it cites.
“Learning Latent Causal Graphs with Unknown Interventions”
Yibo Jiang and Bryon Aragam · 2023
Later among the works it cites.
“Uncovering meanings of embeddings via partial orthogonality”
Yibo Jiang, Bryon Aragam and Victor Veitch · 2023
Later among the works it cites.
“Inference-Time Intervention: Eliciting Truthful Answers from a Language Model”
Kenneth Li et al · 2023
Later among the works it cites.
“Emergent Linear Representations in World Models of Self-Supervised Sequence Models”
Neel Nanda, Andrew Lee and Martin Wattenberg · 2023
Later among the works it cites.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Later among the works it cites.
“The Linear Representation Hypothesis and the Geometry of Large Language Models”, 2023
Kiho Park, Yo Choe and Victor Veitch · 2023
Later among the works it cites.
Goutham Rajendran, Patrik Reizinger, Wieland Brendel and Pradeep Ravikumar · 2023
Later among the works it cites.
“Contrastive Loss is All You Need to Recover Analogies as Parallel Lines”
Narutatsu Ri, Fei-Tzin Lee and Nakul Verma · 2023
Later among the works it cites.
“Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero”
Lisa Schut et al · 2023
Later among the works it cites.
“Linear Representations of Sentiment in Large Language Models”
Curt Tigges, Oskar Hollinsworth, Atticus Geiger and Neel Nanda · 2023
Later among the works it cites.
“Llama: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Later among the works it cites.
“Linear spaces of meanings: compositional structures in vision-language models”
Matthew Trager et al · 2023
Later among the works it cites.
“Score-based Causal Representation Learning with Interventions”
Burak Varici et al · 2023
Later among the works it cites.
“Concept Algebra for Score-based Conditional Model”
Zihao Wang, Lin Gui, Jeffrey Negrea and Victor Veitch · 2023
Later among the works it cites.
“Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models”
Goutham Rajendran et al · 2024
Closest in time.