Fetching the paper…
Reading the bibliography…
The assumption across nearly all language model (LM) tokenization schemes is that tokens should be subwords, i.e., contained within word boundaries.
A new algorithm for data compression
Philip Gage · 1994
Earlier work this paper cites.
Large language models in machine translation
Thorsten Brants, Ashok C. Popat, Peng Xu, Franz J. Och, and Jeffrey Dean · 2007
Earlier work this paper cites.
How many multiword expressions do people know?
Kenneth Church · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S. Gordon · 2011
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2012
Earlier work this paper cites.
The winograd schema challenge
Hector J. Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
Comprehensive annotation of multiword expressions in a social web corpus
Nathan Schneider, Spencer Onuffer, Nora Kazour, Emily Danchik, Michael T. Mordowanec, Henrietta Conrad, and Noah A. Smith · 2014
Earlier work this paper cites.
A word embedding approach to predicting the compositionality of multiword expressions
Bahar Salehi, Paul Cook, and Timothy Baldwin · 2015
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
The indeterminacy of word segmentation and the nature of morphology and syntax
Haspelmath Martin · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Sentencepiece experiments
Taku Kudo · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Earlier work this paper cites.
Investigating the effectiveness of BPE: The power of shorter sequences
Matthias Gallé · 2019
Earlier work this paper cites.
When choosing plausible alternatives, clever hans can be clever
Pride Kavumba, Naoya Inoue, Benjamin Heinzerling, Keshav Singh, Paul Reisert, and Kentaro Inui · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
CoQA: A conversational question answering challenge
Siva Reddy, Danqi Chen, and Christopher D. Manning · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Pre-tokenization of multi-word expressions in cross-lingual word embeddings
Naoki Otani, Satoru Ozaki, Xingyuan Zhao, Yucen Li, Micaelah St Johns, and Lori Levin · 2020
Earlier work this paper cites.
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita · 2020
Earlier work this paper cites.
ProphetNet: Predicting future n-gram for sequence-to-SequencePre-training
Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou · 2020
Earlier work this paper cites.
From english to foreign languages: Transferring pre-trained language models
Ke Tran · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush · 2020
Earlier work this paper cites.
Program synthesis with large language models, 2021
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton · 2021
Earlier work this paper cites.
Evaluating entity disambiguation and the role of popularity in retrieval-based NLP
Anthony Chen, Pallavi Gudipati, Shayne Longpre, Xiao Ling, and Sameer Singh · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Superbizarre is not superb: Derivational morphology improves BERT’s interpretation of complex words
Valentin Hofmann, Janet Pierrehumbert, and Hinrich Schütze · 2021
Cited alongside, same era.
Lattice-BERT: Leveraging multi-granularity representations in Chinese pre-trained language models
Yuxuan Lai, Yijia Liu, Yansong Feng, Songfang Huang, and Dongyan Zhao · 2021
Cited alongside, same era.
Jurassic-1: Technical details and evaluation, 2021
Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham · 2021
Cited alongside, same era.
Bloomberggpt: A large language model for finance, 2023
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Later among the works it cites.
MEGABYTE: Predicting million-byte sequences with multiscale transformers
Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis · 2023
Later among the works it cites.
MAGNET: Improving the multilingual fairness of language models with adaptive gradient-based tokenization
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Valentin Hofmann, Tomasz Limisiewicz, Yulia Tsvetkov, and Noah A. Smith · 2024
Later among the works it cites.
Getting the most out of your tokenizer for pre-training and domain adaptation
Gautier Dagan, Gabriel Synnaeve, and Baptiste Rozière · 2024
Later among the works it cites.
CUTE: Measuring LLMs’ understanding of their tokens
Lukas Edman, Helmut Schmid, and Alexander Fraser · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y. Lee, Benoît Sagot, and Samson Tan · 2021
Cited alongside, same era.
Investigating the limitations of transformers with simple arithmetic tasks, 2021
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin · 2021
Cited alongside, same era.
Winogrande: an adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2021
Cited alongside, same era.
Scale efficiently: Insights from pre-training and fine-tuning transformers
Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler · 2021
Cited alongside, same era.
Representing numbers in NLP: a survey and a vision
Avijit Thawani, Jay Pujara, Filip Ilievski, and Pedro Szekely · 2021
Cited alongside, same era.
AMBERT: A pre-trained language model with multi-grained tokenization
Xinsong Zhang, Pengshuai Li, and Hang Li · 2021
Cited alongside, same era.
Canine: Pre-training an efficient tokenization-free encoder for language representation
Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting · 2022
Cited alongside, same era.
Better & faster large language models via multi-token prediction
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Roziere, David Lopez-Paz, and Gabriel Synnaeve · 2024
Later among the works it cites.
Unpacking tokenization: Evaluating text compression and its correlation with model performance
Omer Goldman, Avi Caciularu, Matan Eyal, Kris Cao, Idan Szpektor, and Reut Tsarfaty · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology, 2024
Google · 2024
Later among the works it cites.
Data mixture inference: What do BPE tokenizers reveal about their training data?
Jonathan Hayase, Alisa Liu, Yejin Choi, Sewoong Oh, and Noah A. Smith · 2024
Later among the works it cites.
The remarkable robustness of llms: Stages of inference?, 2024
Vedang Lad, Wes Gurnee, and Max Tegmark · 2024
Later among the works it cites.
A short introduction to pre-tokenization weirdness, 2024
Sander Land · 2024
Later among the works it cites.
Fishing for magikarp: Automatically detecting under-trained tokens in large language models
Sander Land and Max Bartolo · 2024
Later among the works it cites.
Datacomp-lm: In search of the next generation of training sets for language models, 2024
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan Bitton, Marianna Nezhurina, Amro Abbas, Cheng-Yu Hsieh, Dhruba Ghosh, Josh Gardner, Maciej Kilian, Hanlin Zhang, Rulin Shao, Sarah Pratt, Sunny Sanyal, Gabriel Ilharco, Giannis Daras, Kalyani Marathe, Aaron Gokaslan, Jieyu Zhang, Khyathi Chandu, Thao Nguyen, Igor Vasiljevic, Sham Kakade, Shuran Song, Sujay Sanghavi, Fartash Faghri, Sewoong Oh, Luke Zettlemoyer, Kyle Lo, Alaaeldin El-Nouby, Hadi Pouransari, Alexander Toshev, Stephanie Wang, Dirk Groeneveld, Luca Soldaini, Pang Wei Koh, Jenia Jitsev, Thomas Kollar, Alexandros G. Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar · 2024
Later among the works it cites.
Ofa: A framework of initializing unseen subword embeddings for efficient large-scale multilingual continued pretraining
Yihong Liu, Peiqin Lin, Mingyang Wang, and Hinrich Schütze · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
Meta · 2024
Later among the works it cites.
Zero-shot tokenizer transfer
Benjamin Minixhofer, Edoardo Ponti, and Ivan Vulić · 2024
Later among the works it cites.
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, Nathan Lambert, Dustin Schwenk, Oyvind Tafjord, Taira Anderson, David Atkinson, Faeze Brahman, Christopher Clark, Pradeep Dasigi, Nouha Dziri, Michal Guerquin, Hamish Ivison, Pang Wei Koh, Jiacheng Liu, Saumya Malik, William Merrill, Lester James V. Miranda, Jacob Morrison, Tyler Murray, Crystal Nam, Valentina Pyatkin, Aman Rangapur, Michael Schmitz, Sam Skjonsberg, David Wadden, Christopher Wilhelm, Michael Wilson, Luke Zettlemoyer, Ali Farhadi, Noah A. Smith, and Hannaneh Hajishirzi · 2024
Later among the works it cites.
Hello GPT-4o, 2024
OpenAI · 2024
Later among the works it cites.
Byte latent transformer: Patches scale better than tokens, 2024
Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, and Srinivasan Iyer · 2024
Later among the works it cites.
Understanding and mitigating tokenization bias in language models, 2024
Buu Phan, Marton Havasi, Matthew Muckley, and Karen Ullrich · 2024
Later among the works it cites.
Tokenization is more than compression
Craig W Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner · 2024
Later among the works it cites.
Tokenization counts: the impact of tokenization on arithmetic in frontier llms, 2024
Aaditya K. Singh and DJ Strouse · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
Scaling laws with vocabulary: Larger models deserve larger vocabularies
Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin, and Ngai Wong · 2024
Later among the works it cites.
From language models over tokens to language models over characters
Tim Vieira, Ben LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Brian DuSell, John Terilla, Timothy J O’Donnell, and Ryan Cotterell · 2024
Later among the works it cites.
Mambabyte: Token-free selective state space model
Junxiong Wang, Tushaar Gangavarapu, Jing Nathan Yan, and Alexander M Rush · 2024
Later among the works it cites.
AGIEval: A human-centric benchmark for evaluating foundation models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan · 2024
Later among the works it cites.
Deepseek-v3 technical report, 2025
DeepSeek-AI · 2025
Closest in time.
Over-tokenized transformer: Vocabulary is generally worth scaling, 2025
Hongzhi Huang, Defa Zhu, Banggu Wu, Yutao Zeng, Ya Wang, Qiyang Min, and Xun Zhou · 2025
Closest in time.
Dynamic chunking for end-to-end hierarchical sequence modeling, 2025
Sukjun Hwang, Brandon Wang, and Albert Gu · 2025
Closest in time.
From tokens to words: On the inner lexicon of LLMs
Guy Kaplan, Matanel Oren, Yuval Reif, and Roy Schwartz · 2025
Closest in time.
Universal cross-tokenizer distillation via approximate likelihood matching
Benjamin Minixhofer, Ivan Vulić, and Edoardo Maria Ponti · 2025
Closest in time.
Stochastok: Improving fine-grained subword understanding in LLMs
Anya Sims, Cong Lu, Klara Kaleb, Jakob Nicolaus Foerster, and Yee Whye Teh · 2025
Closest in time.
Egalitarian language representation in language models: It all begins with tokenizers
Menan Velayuthan and Kengatharaiyer Sarveswaran · 2025
Closest in time.