Fetching the paper…
Reading the bibliography…
Pre-tokenization, the initial step in many modern tokenization pipelines, segments text into smaller units called pretokens, typically splitting on whitespace and punctuation.
Transmission of information, 1961
Robert M Fano and WT Wintringham · 1961
Earlier work this paper cites.
Word association norms, mutual information, and lexicography
Kenneth Ward Church and Patrick Hanks · 1989
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage · 1994
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima · 2012
Earlier work this paper cites.
A word embedding approach to predicting the compositionality of multiword expressions
Bahar Salehi, Paul Cook, and Timothy Baldwin · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo · 2018
Earlier work this paper cites.
Investigating the effectiveness of BPE: The power of shorter sequences
Matthias Gallé · 2019
Earlier work this paper cites.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett · 2020
Earlier work this paper cites.
Pre-tokenization of multi-word expressions in cross-lingual word embeddings
Naoki Otani, Satoru Ozaki, Xingyuan Zhao, Yucen Li, Micaelah St Johns, and Lori Levin · 2020
Earlier work this paper cites.
BPE beyond word boundary: How NOT to use multi word expressions in neural machine translation
Dipesh Kumar and Avijit Thawani · 2022
Cited alongside, same era.
Rare tokens degenerate all tokens: Improving neural text generation via adaptive gradient gating for rare token embeddings
Sangwon Yu, Jongyoon Song, Heeseung Kim, Seongmin Lee, Woo-Jong Ryu, and Sungroh Yoon · 2022
Cited alongside, same era.
Multi-word tokenization for sequence compression
Leonidas Gee, Leonardo Rigutini, Marco Ernandes, and Andrea Zugarini · 2023
Cited alongside, same era.
The minipile challenge for data-efficient language models, 2023
Jean Kaddour · 2023
Cited alongside, same era.
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell · 2023
Cited alongside, same era.
Unpacking tokenization: Evaluating text compression and its correlation with model performance
Omer Goldman, Avi Caciularu, Matan Eyal, Kris Cao, Idan Szpektor, and Reut Tsarfaty · 2024
Later among the works it cites.
Haoran Lian, Yizhe Xiong, Jianwei Niu, Shasha Mo, Zhenpeng Su, Zijia Lin, Hui Chen, Peng Liu, Jungong Han, and Guiguang Ding · 2024
Later among the works it cites.
Tokenization is more than compression
Craig W Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner · 2024
Later among the works it cites.
Tokenization counts: the impact of tokenization on arithmetic in frontier llms, 2024
Aaditya K. Singh and DJ Strouse · 2024
Later among the works it cites.
MiLe loss: a new loss for mitigating the bias of learning difficulties in generative language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tokenizer choice for LLM training: Negligible or crucial?
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Buschhoff, Charvi Jain, Alexander Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suarez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, and Nicolas Flores-Herr · 2024
Cited alongside, same era.
BPE-knockout: Pruning pre-existing BPE tokenisers with backwards-compatible morphological semi-supervision
Thomas Bauwens and Pieter Delobelle · 2024
Cited alongside, same era.
BPE gets picky: Efficient vocabulary refinement during tokenizer training
Pavel Chizhov, Catherine Arnett, Elizaveta Korotkova, and Ivan P. Yamshchikov · 2024
Cited alongside, same era.
An analysis of BPE vocabulary trimming in neural machine translation
Marco Cognetta, Tatsuya Hiraoka, Rico Sennrich, Yuval Pinter, and Naoaki Okazaki · 2024
Cited alongside, same era.
Two counterexamples to tokenization and the noiseless channel
Marco Cognetta, Vilém Zouhar, Sangwhan Moon, and Naoaki Okazaki · 2024
Cited alongside, same era.
Getting the most out of your tokenizer for pre-training and domain adaptation
Gautier Dagan, Gabriel Synnaeve, and Baptiste Roziere · 2024
Cited alongside, same era.
Zhenpeng Su, Zijia Lin, Baixue Baixue, Hui Chen, Songlin Hu, Wei Zhou, Guiguang Ding, and Xing W · 2024
Later among the works it cites.
Don’t touch my diacritics, 2025
Kyle Gorman and Yuval Pinter · 2025
Closest in time.
Over-tokenized transformer: Vocabulary is generally worth scaling, 2025
Hongzhi Huang, Defa Zhu, Banggu Wu, Yutao Zeng, Ya Wang, Qiyang Min, and Xun Zhou · 2025
Closest in time.
Superbpe: Space travel for language models, 2025
Alisa Liu, Jonathan Hayase, Valentin Hofmann, Sewoong Oh, Noah A. Smith, and Yejin Choi · 2025
Closest in time.
How much is enough? the diminishing returns of tokenization training data
Varshini Reddy, Craig W Schmidt, Yuval Pinter, and Chris Tanner · 2025
Closest in time.
Egalitarian language representation in language models: It all begins with tokenizers
Menan Velayuthan and Kengatharaiyer Sarveswaran · 2025
Closest in time.
Tokenization is sensitive to language variation, 2025
Anna Wegmann, Dong Nguyen, and David Jurgens · 2025
Closest in time.