Fetching the paper…
Reading the bibliography…
We present a training-free method to transplant tokenizers in pretrained large language models (LLMs) by reconstructing unseen token embeddings via Orthogonal Matching Pursuit (OMP).
Orthogonal matching pursuit: Recursive function approximation with applications to wavelet decomposition
Yagyensh Chandra Pati, Ramin Rezaiifar, and Perinkulam Sambamurthy Krishnaprasad · 1993
Earlier work this paper cites.
Signal recovery from random measurements via orthogonal matching pursuit
Joel A Tropp and Anna C Gilbert · 2007
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Clustering of words using dictionary-learnt word representations
Remya R. K. Menon, S Gargi, and S Samili · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Orthogonal matching pursuit for text classification
Konstantinos Skianis, Nikolaos Tziortziotis, and Michalis Vazirgiannis · 2018
Earlier work this paper cites.
How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings
Kawin Ethayarajh · 2019
Earlier work this paper cites.
Paws-x: A cross-lingual adversarial dataset for paraphrase identification
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge · 2019
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
Cross-lingual alignment methods for multilingual BERT: A comparative study
Saurabh Kulshreshtha, Jose Luis Redondo Garcia, and Ching-Yun Chang · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Cited alongside, same era.
Initializing new word embeddings for pretrained language models, 2021
John Hewitt · 2021
Efficient language model training through cross-lingual and progressive transfer learning, 2023
Malte Ostendorff and Georg Rehm · 2023
Later among the works it cites.
Hydra: Sequentially-dependent draft heads for medusa decoding
Zachary Ankner, Rishab Parthasarathy, Aniruddha Nrusimha, Christopher Rinard, Jonathan Ragan-Kelley, and William Brandon · 2024
Later among the works it cites.
Towards cross-tokenizer distillation: the universal logit distillation loss for llms
Nicolas Boizard, Kevin El Haddad, Céline Hudelot, and Pierre Colombo · 2024
Later among the works it cites.
Arcee‘s MergeKit: A toolkit for merging large language models
Charles Goddard, Shamane Siriwardhana, Malikeh Ehghaghi, Luke Meyers, Vladimir Karpukhin, Brian Benedict, Mark McQuade, and Jacob Solawetz · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych · 2021
Cited alongside, same era.
Fast vocabulary transfer for language model compression
Leonidas Gee, Andrea Zugarini, Leonardo Rigutini, and Paolo Torroni · 2022
Cited alongside, same era.
WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models
Benjamin Minixhofer, Fabian Paischer, and Navid Rekabsaz · 2022
Cited alongside, same era.
FOCUS: Effective embedding initialization for monolingual specialization of multilingual models
Konstantin Dobler and Gerard de Melo · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Cited alongside, same era.
Word translation without parallel data, 2018a
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou
Cited in the paper.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Later among the works it cites.
Zero-shot tokenizer transfer, 2024
Benjamin Minixhofer, Edoardo Maria Ponti, and Ivan Vulić · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Vocab transplantation tool (github repository), 2025
jukofyork · 2025
Closest in time.
Language models use trigonometry to do addition
Subhash Kantamneni and Max Tegmark · 2025
Closest in time.
Shared global and local geometry of language model embeddings, 2025
Andrew Lee, Melanie Weber, Fernanda Viégas, and Martin Wattenberg · 2025
Closest in time.