Fetching the paper…
Reading the bibliography…
Pretrained language models (LLMs) are often constrained by their fixed tokenization schemes, leading to inefficiencies and performance limitations, particularly for multilingual or specialized applications.
Don’t stop pretraining: Adapt language models to domains and tasks, 2020
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith · 2004
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
As good as new. how to successfully recycle English GPT-2 to make models for other languages
Wietse de Vries, Andreas van Cranenburgh, and Malvina Nissim · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021
Jesse Dodge, Maarten Sap, Ana Marasovi, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Earlier work this paper cites.
Oscar Minixhofer, Fabian Paischer, and Navid Rekabsaz · 2021
Earlier work this paper cites.
MegaByte: Predicting million-byte sequences with multiscale transformers
Lasha Ahia, Alexander Poli, Christopher Ré, Matei Zaharia, and Stefano Ermon · 2023
Earlier work this paper cites.
Tamil-llama: A new tamil language model based on llama 2, 2023
Abhinand Balachandran and Arun Raj M · 2023
Cited alongside, same era.
Attention is all you need, 2023
JAshish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2023
Cited alongside, same era.
Cost effective continual pre-training for bridging the llm performance gap for indic languages, 2024
Vamshi Krishna Aribandi, Arpan Mandal, and Nikesh Garera · 2024
Cited alongside, same era.
Retok: Replacing tokenizer to enhance representation efficiency in large language model, 2024
Zhen Chen, Jianing Wang, Qiushi Sun, Xiubo Geng, Nuo Xu, Wenji Mao, and Daxin Jiang · 2024
Cited alongside, same era.
Andrew R. Gee and Christopher D. Manning · 2024
Later among the works it cites.
Zero-shot tokenizer transfer, 2024
Oscar Minixhofer, Marcelo Orenes-Vera, and Ivan Vulić · 2024
Later among the works it cites.
Llamaturk: Adapting open-source generative large language models for low-resource language, 2024
Emincan Uygun, Erion Çano, and İzzet Emre Kiciman · 2024
Later among the works it cites.
Automathtext: Autonomous data selection with language models for mathematical texts, 2024
Yifan Zhang, Yifan Luo, Yang Yuan, and Andrew Chi-Chih Yao · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Miriam Dobler and Gerard de Melo · 2024
Cited alongside, same era.
Airavata: Introducing Hindi Instruction-tuned LLM, 2024
Jay Gala, Thanmay Jayakumar, Jaavid Aktar Husain, Aswanth Kumar M, Mohammed Safi Ur Rahman Khan, Diptesh Kanojia, Ratish Puduppully, Mitesh M. Khapra, Raj Dabre, Rudra Murthy, and Anoop Kunchukuttan · 2024
Cited alongside, same era.
github-code
codeparrot
Cited in the paper.
meta-llama/llama-3.2-3b
Meta-Llama
Cited in the paper.
Focus: Effective embedding initialization for language adaptation of large language models, 2023a
Oscar Minixhofer, Gábor Berend, and Jie Yang
Cited in the paper.
Efficient language model training through cross-lingual and progressive transfer learning, 2023b
Oscar Minixhofer, Fabian Paischer, and Navid Rekabsaz
Cited in the paper.
Qwen/qwen2.5-3b
Qwen
Cited in the paper.
mc4-hindi-cleaned-3.0
zicsx
Cited in the paper.
Alisa Liu, Sang Michael Xie, Michail Papauschek, Graham Neubig, Jonathan Frankle, Volodymyr Kuleshov, and Ankit Singh · 2025
Closest in time.