Fetching the paper…
Reading the bibliography…
Modern language models can process inputs across diverse languages and modalities.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 1910
Earlier work this paper cites.
The case for case
Charles J. Fillmore · 1968
Earlier work this paper cites.
Semantic Interpretation in Generative Grammar
R.S. Jackendoff · 1974
Earlier work this paper cites.
Observations on Number Systems and Semantics
Donald C Laycock · 1975
Earlier work this paper cites.
Events in the Semantics of English: A Study in Subatomic Semantics
Terence Parsons · 1990
Earlier work this paper cites.
When is ”nearest neighbor” meaningful?
Kevin S. Beyer, Jonathan Goldstein, Raghu Ramakrishnan, and Uri Shaft · 1999
Earlier work this paper cites.
Where do you know what you know? the representation of semantic knowledge in the human brain
Karalyn Patterson, Peter J Nestor, and Timothy T Rogers · 2007
Earlier work this paper cites.
Exploiting similarities among languages for machine translation
Tomas Mikolov, Quoc V Le, and Ilya Sutskever · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and Larry Zitnick · 2014
Earlier work this paper cites.
Gale phase 3 and 4 chinese newswire parallel text
Song Chen, Gary Krug, and Stephanie Strassel · 2016
Earlier work this paper cites.
The hub-and-spoke hypothesis of semantic memory
Karalyn Patterson and Matthew A Lambon Ralph · 2016
Earlier work this paper cites.
Learning bilingual word embeddings with (almost) no bilingual data
Mikel Artetxe, Gorka Labaka, and Eneko Agirre · 2017
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Matthew Honnibal and Ines Montani · 2017
Earlier work this paper cites.
The neural and computational bases of semantic cognition
Matthew A Lambon Ralph, Elizabeth Jefferies, Karalyn Patterson, and Timothy T Rogers · 2017
Earlier work this paper cites.
Offline bilingual word vectors, orthogonal transformations and the inverted softmax
Samuel L. Smith, David H. P. Turban, Steven Hamblin, and Nils Y. Hammerla · 2017
Earlier work this paper cites.
Word translation without parallel data
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou · 2018
Earlier work this paper cites.
The democratization of deep learning in tass 2017
Manuel Carlos Díaz-Galiano, Eugenio Martínez-Cámara, Miguel Ángel García Cumbreras, Manuel García Vega, and Julio Villena-Román · 2018
Earlier work this paper cites.
Cross-lingual alignment of contextual word embeddings, with applications to zero-shot dependency parsing
Tal Schuster, Ori Ram, Regina Barzilay, and Amir Globerson · 2019
Earlier work this paper cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
Multilingual alignment of contextual word representations
Steven Cao, Nikita Kitaev, and Dan Klein · 2020
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Earlier work this paper cites.
The multilingual Amazon reviews corpus
Phillip Keung, Yichao Lu, György Szarvas, and Noah A. Smith · 2020
Earlier work this paper cites.
COGS: A compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen · 2020
Earlier work this paper cites.
Interpreting GPT: the logit lens
nostalgebraist · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2021
Earlier work this paper cites.
Probing contextual language models for common ground with visual representations
Gabriel Ilharco, Rowan Zellers, Ali Farhadi, and Hannaneh Hajishirzi · 2021
Cited alongside, same era.
Implicit representations of meaning in neural language models
Belinda Z. Li, Maxwell Nye, and Jacob Andreas · 2021
Cited alongside, same era.
Pretrained transformers as universal computation engines
Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy · 2021
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Gemini-Team · 2024
Closest in time.
Patchscopes: A unifying framework for inspecting hidden representations of language models
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva · 2024
Closest in time.
Listen, think, and understand
Yuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky, and James R. Glass · 2024
Closest in time.
The unreasonable ineffectiveness of the deeper layers
Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, and Daniel A. Roberts · 2024
Closest in time.
mOthello: When do cross-lingual representation alignment and cross-lingual transfer emerge in multilingual models?
Tianze Hua, Tian Yun, and Ellie Pavlick · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Infusing Finetuning with Semantic Dependencies
Zhaofeng Wu, Hao Peng, and Noah A. Smith · 2021
Cited alongside, same era.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel · 2021
Cited alongside, same era.
Causal scrubbing, a method for rigorously testing interpretability hypotheses
Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldwosky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas · 2022
Cited alongside, same era.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2022
Cited alongside, same era.
Linearly mapping from image to text space
Jack Merullo, Louis Castricato, Carsten Eickhoff, and Ellie Pavlick · 2022
Cited alongside, same era.
Extracting latent steering vectors from pretrained language models
Nishant Subramani, Nivedita Suresh, and Matthew Peters · 2022
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Cited alongside, same era.
The platonic representation hypothesis
Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola · 2024
Closest in time.
Exploring concept depth: How large language models acquire knowledge at different layers?
Mingyu Jin, Qinkai Yu, Jingyuan Huang, Qingcheng Zeng, Zhenting Wang, Wenyue Hua, Haiyan Zhao, Kai Mei, Yanda Meng, Kaize Ding, et al · 2024
Closest in time.
The Llama-3-Team · 2024
Closest in time.
Unified-IO 2: Scaling autoregressive multimodal models with vision language audio and action
Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, and Aniruddha Kembhavi · 2024
Closest in time.
Task vectors are cross-modal, 2024
Grace Luo, Trevor Darrell, and Amir Bar · 2024
Closest in time.
Do vision and language encoders represent the world similarly?
Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Mohamed El Amine Seddik, Sanath Narayan, Karttikeya Mangalam, and Noel E. O’Connor · 2024
Closest in time.
MDBG Chinese-English dictionary (CC-CEDICT)
MDBG · 2024
Closest in time.
Language models implement simple Word2Vec-style vector arithmetic
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2024
Closest in time.
What do language models hear? probing for auditory representations in language models
Jerry Ngo and Yoon Kim · 2024
Closest in time.
Steering llama 2 via contrastive activation addition
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner · 2024
Closest in time.
Pre-training small base lms with fewer tokens
Sunny Sanyal, Sujay Sanghavi, and Alexandros G. Dimakis · 2024
Closest in time.
Jieba: Chinese text segmentation tool
Junyi Sun · 2024
Closest in time.
Language-specific neurons: The key to multilingual capabilities in large language models
Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.
Function vectors in large language models
Eric Todd, Millicent Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau · 2024
Closest in time.
Diffusion lens: Interpreting text encoders in text-to-image pipelines
Michael Toker, Hadas Orgad, Mor Ventura, Dana Arad, and Yonatan Belinkov · 2024
Closest in time.
Activation addition: Steering language models without optimization
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid · 2024
Closest in time.
Multilingual e5 text embeddings: A technical report
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei · 2024
Closest in time.
Do llamas work in English? on the latent language of multilingual transformers
Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West · 2024
Closest in time.
pyvene: A library for understanding and improving pytorch models via interventions, 2024
Zhengxuan Wu, Atticus Geiger, Aryaman Arora, Jing Huang, Zheng Wang, Noah D. Goodman, Christopher D. Manning, and Christopher Potts · 2024
Closest in time.
Do large language models latently perform multi-hop reasoning?
Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, and Sebastian Riedel · 2024
Closest in time.
Hongchuan Zeng, Senyu Han, Lu Chen, and Kai Yu · 2024
Closest in time.
How do large language models handle multilingualism?
Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi, and Lidong Bing · 2024
Closest in time.