Fetching the paper…
Reading the bibliography…
Tokenization, a crucial initial step in natural language processing, is governed by several key parameters, such as the tokenization algorithm, vocabulary size, pre-tokenization strategy, inference strategy, and training data corpus.
Human Behavior and the Principle of Least Effort
Zipf, G. K · 1949
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Japanese and korean voice search
Schuster, M. and Nakajima, K · 2012
Earlier work this paper cites.
A word embedding approach to predicting the compositionality of multiword expressions
Salehi, B., Cook, P., and Baldwin, T · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Kudo, T · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2020
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Finding the optimal vocabulary size for neural machine translation
Gowda, T. and May, J · 2020
Earlier work this paper cites.
Mielke, S. J., Alyafeai, Z., Salesky, E., Raffel, C., Dey, M., Gallé, M., Raja, A., Si, C., Lee, W. Y., Sagot, B., and Tan, S · 2021
Earlier work this paper cites.
Study of various methods for tokenization
Rai, A. and Borah, S · 2021
Cited alongside, same era.
How good is your tokenizer? on the monolingual performance of multilingual language models
Rust, P., Pfeiffer, J., Vulić, I., Ruder, S., and Gurevych, I · 2021
Cited alongside, same era.
Training compute-optimal large language models, 2022
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Cited alongside, same era.
BPE beyond word boundary: How NOT to use multi word expressions in neural machine translation
Kumar, D. and Thawani, A · 2022
Cited alongside, same era.
Analyzing cognitive plausibility of subword tokenization
Beinborn, L. and Pinter, Y · 2023
Cited alongside, same era.
OSCaR: Object state captioning and state change representation
Nguyen, N., Bi, J., Vosoughi, A., Tian, Y., Fazli, P., and Xu, C · 2024
Later among the works it cites.
Scaling laws for pre-training agents and world models, 2024
Pearce, T., Rashid, T., Bignell, D., Georgescu, R., Devlin, S., and Hofmann, K · 2024
Later among the works it cites.
Tokenization is more than compression
Schmidt, C. W., Reddy, V., Zhang, H., Alameddine, A., Uzan, O., Pinter, Y., and Tanner, C · 2024
Later among the works it cites.
Greed is all you need: An evaluation of tokenizer inference methods
Uzan, O., Schmidt, C. W., Tanner, C., and Pinter, Y · 2024
Later among the works it cites.
Redpajama: an open dataset for training large language models, 2024
Weber, M., Fu, D., Anthony, Q., Oren, Y., Adams, S., Alexandrov, A., Lyu, X., Nguyen, H., Yao, X., Adams, V., Athiwaratkun, B., Chalamala, R., Chen, K., Ryabinin, M., Dao, T., Liang, P., Ré, C., Rish, I., and Zhang, C · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tokenization impacts multilingual language modeling: Assessing vocabulary allocation and overlap across languages
Limisiewicz, T., Balhar, J., and Mareček, D · 2023
Cited alongside, same era.
Incorporating context into subword vocabularies
Yehezkel, S. and Pinter, Y · 2023
Cited alongside, same era.
Tokenization and the noiseless channel
Zouhar, V., Meister, C., Gastaldi, J., Du, L., Sachan, M., and Cotterell, R · 2023
Cited alongside, same era.
Tokenizer choice for LLM training: Negligible or crucial?
Ali, M., Fromm, M., Thellmann, K., Rutmann, R., Lübbering, M., Leveling, J., Klug, K., Ebert, J., Doll, N., Buschhoff, J., Jain, C., Weber, A., Jurkschat, L., Abdelwahab, H., John, C., Ortiz Suarez, P., Ostendorff, M., Weinbach, S., Sifa, R., Kesselheim, S., and Flores-Herr, N · 2024
Cited alongside, same era.
Getting the most out of your tokenizer for pre-training and domain adaptation, 2024
Dagan, G., Synnaeve, G., and Rozière, B · 2024
Cited alongside, same era.
Later among the works it cites.
When scaling meets LLM finetuning: The effect of data, model and finetuning method
Zhang, B., Liu, Z., Cherry, C., and Firat, O · 2024
Later among the works it cites.
Superbpe: Space travel for language models, 2025
Liu, A., Hayase, J., Hofmann, V., Oh, S., Smith, N. A., and Choi, Y · 2025
Closest in time.
Boundless byte pair encoding: Breaking the pre-tokenization barrier, 2025
Schmidt, C. W., Reddy, V., Tanner, C., and Pinter, Y · 2025
Closest in time.
Egalitarian language representation in language models: It all begins with tokenizers
Velayuthan, M. and Sarveswaran, K · 2025
Closest in time.