Fetching the paper…
Reading the bibliography…
Informally, the 'linear representation hypothesis' is the idea that high-level concepts are represented linearly as directions in some representation space.
Linguistic regularities in continuous space word representations
Mikolov, T., Yih, W.-T., and Zweig, G · 2013
Earlier work this paper cites.
word2vec explained: deriving Mikolov et al.’s negative-sampling word-embedding method
Goldberg, Y. and Levy, O · 2014
Earlier work this paper cites.
Linguistic regularities in sparse and explicit word representations
Levy, O. and Goldberg, Y · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
A latent variable model approach to PMI-based word embeddings
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A · 2016
Earlier work this paper cites.
Generating sentences from a continuous space
Bowman, S. R., Vilnis, L., Vinyals, O., Dai, A., Jozefowicz, R., and Bengio, S · 2016
Earlier work this paper cites.
Word embeddings, analogies, and machine learning: Beyond king - man + woman = queen
Drozd, A., Gladkova, A., and Matsuoka, S · 2016
Earlier work this paper cites.
Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn’t
Gladkova, A., Drozd, A., and Matsuoka, S · 2016
Earlier work this paper cites.
beta-VAE: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2016
Earlier work this paper cites.
Unsupervised feature extraction by time-contrastive learning and nonlinear ICA
Hyvarinen, A. and Morioka, H · 2016
Earlier work this paper cites.
Take and took, gaggle and goose, book and read: Evaluating the utility of vector differences for lexical relation learning
Vylomova, E., Rimell, L., Cohn, T., and Baldwin, T · 2016
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y · 2017
Earlier work this paper cites.
The strange geometry of skip-gram with negative sampling
Mimno, D. and Thompson, L · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Towards a definition of disentangled representations
Higgins, I., Amos, D., Pfau, D., Racaniere, S., Matthey, L., Rezende, D., and Lerchner, A · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Word translation without parallel data
Lample, G., Conneau, A., Ranzato, M., Denoyer, L., and Jégou, H · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings
Ethayarajh, K · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
Hewitt, J. and Manning, C. D · 2019
Cited alongside, same era.
Visualizing and measuring the geometry of BERT
Reif, E., Yuan, A., Wattenberg, M., Viegas, F. B., Coenen, A., Pearce, A., and Kim, B · 2019
Cited alongside, same era.
A survey of cross-lingual word embedding models
Ruder, S., Vulić, I., and Søgaard, A · 2019
Cited alongside, same era.
Understanding the source of semantic regularities in word embeddings
Chiang, H.-Y., Camacho-Collados, J., and Pardos, Z · 2020
Cited alongside, same era.
word2word: A collection of bilingual lexicons for 3,564 language pairs
Choe, Y. J., Park, K., and Kim, D · 2020
Cited alongside, same era.
Analogies minus analogy test: measuring regularities in word embeddings
Fournier, L., Dupoux, E., and Dunbar, E · 2020
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Geva, M., Caciularu, A., Wang, K., and Goldberg, Y · 2022
Later among the works it cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M · 2022
Later among the works it cites.
Locating and editing factual associations in GPT
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Later among the works it cites.
Understanding linearity of cross-lingual word embedding mappings
Peng, X., Stevenson, M., Lin, C., and Li, C · 2022
Later among the works it cites.
Language models represent space and time
Gurnee, W. and Tegmark, M · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Variational autoencoders and nonlinear ICA: A unifying framework
Khemakhem, I., Kingma, D., Monti, R., and Hyvarinen, A · 2020
Cited alongside, same era.
On the sentence embeddings from pre-trained language models
Li, B., Zhou, H., He, J., Wang, M., Yang, Y., and Li, L · 2020
Cited alongside, same era.
Interpreting GPT: the logit lens, 2020
nostalgebraist · 2020
Cited alongside, same era.
Sentence analogies: Linguistic regularities in sentence embeddings
Zhu, X. and de Melo, G · 2020
Cited alongside, same era.
Probing BERT in hyperbolic spaces
Chen, B., Fu, Y., Xu, G., Xie, P., Tan, C., Chen, M., and Jing, L · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al · 2021
Cited alongside, same era.
Hendel, R., Geva, M., and Globerson, A · 2023
Closest in time.
Linearity of relation decoding in transformer language models
Hernandez, E., Sharma, A. S., Haklay, T., Meng, K., Wattenberg, M., Andreas, J., Belinkov, Y., and Bau, D · 2023
Closest in time.
Uncovering meanings of embeddings via partial orthogonality
Jiang, Y., Aragam, B., and Veitch, V · 2023
Closest in time.
Language models implement simple word2vec-style vector arithmetic
Merullo, J., Eickhoff, C., and Pavlick, E · 2023
Closest in time.
Emergent linear representations in world models of self-supervised sequence models
Nanda, N., Lee, A., and Wattenberg, M · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Prompt algebra for task composition
Perera, P., Trager, M., Zancato, L., Achille, A., and Soatto, S · 2023
Closest in time.
Function vectors in large language models
Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Closest in time.
Linear spaces of meanings: Compositional structures in vision-language models
Trager, M., Perera, P., Zancato, L., Achille, A., Bhatia, P., and Soatto, S · 2023
Closest in time.
Activation addition: Steering language models without optimization
Turner, A. M., Thiergart, L., Udell, D., Leech, G., Mini, U., and MacDiarmid, M · 2023
Closest in time.
Concept algebra for score-based conditional models
Wang, Z., Gui, L., Negrea, J., and Veitch, V · 2023
Closest in time.
Representation engineering: A top-down approach to AI transparency
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., Goel, S., Li, N., Byun, M. J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., Kolter, Z., and Hendrycks, D · 2023
Closest in time.
Gemma: Open models based on gemini research and technology
Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Closest in time.