Fetching the paper…
Reading the bibliography…
State-of-the-art language models are autoregressive and operate on subword units known as tokens.
Data compression using adaptive coding and partial string matching
Cleary, J. and Witten, I · 1984
Earlier work this paper cites.
The context-tree weighting method: Basic properties
Willems, F. M., Shtarkov, Y. M., and Tjalkens, T. J · 1995
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Investigating the effectiveness of bpe: The power of shorter sequences
Gallé, M · 2019
Earlier work this paper cites.
Bpe-dropout: Simple and effective subword regularization
Provilkov, I., Emelianenko, D., and Voita, E · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Song, X., Salcianu, A., Song, Y., Dopson, D., and Zhou, D · 2020
Earlier work this paper cites.
You should evaluate your language model on marginal likelihood over tokenisations
Cao, K. and Rimell, L · 2021
Earlier work this paper cites.
Charformer: Fast character transformers via gradient-based subword tokenization
Tay, Y., Tran, V. Q., Ruder, S., Gupta, J., Chung, H. W., Bahri, D., Qin, Z., Baumgartner, S., Yu, C., and Metzler, D · 2021
Earlier work this paper cites.
Efficient transformers with dynamic token pooling
Nawrot, P., Chorowski, J., Łańcucki, A., and Ponti, E. M · 2022
Cited alongside, same era.
Byt5: Towards a token-free future with pre-trained byte-to-byte models
Xue, L., Barua, A., Constant, N., Al-Rfou, R., Narang, S., Kale, M., Roberts, A., and Raffel, C · 2022
Cited alongside, same era.
Improving language plasticity via pretraining with active forgetting
Chen, Y., Marchisio, K., Raileanu, R., Adelani, D., Saito Stenetorp, P. L. E., Riedel, S., and Artetxe, M · 2023
Cited alongside, same era.
Should you marginalize over possible tokenizations?
Chirkova, N., Kruszewski, G., Rozen, J., and Dymetman, M · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini, T · 2023
Cited alongside, same era.
Two counterexamples to \ \backslash textit { \{ Tokenization and the Noiseless Channel } \}
Cognetta, M., Zouhar, V., Moon, S., and Okazaki, N · 2024
Closest in time.
Getting the most out of your tokenizer for pre-training and domain adaptation
Dagan, G., Synnaeve, G., and Rozière, B · 2024
Closest in time.
Unpacking tokenization: Evaluating text compression and its correlation with model performance
Goldman, O., Caciularu, A., Eyal, M., Cao, K., Szpektor, I., and Tsarfaty, R · 2024
Closest in time.
Attention with markov: A framework for principled analysis of transformers via markov chains
Makkuva, A. V., Bondaschi, M., Girish, A., Nagle, A., Jaggi, M., Kim, H., and Gastpar, M · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Think before you speak: Training language models with pause tokens
Goyal, S., Ji, Z., Rawat, A. S., Menon, A. K., Kumar, S., and Nagarajan, V · 2023
Cited alongside, same era.
Guidance ai, 2023
guidance ai · 2023
Cited alongside, same era.
Languages through the looking glass of bpe compression
Gutierrez-Vasques, X., Bentz, C., and Samardžić, T · 2023
Cited alongside, same era.
Task-adaptive tokenization: Enhancing long-form text generation efficacy in mental health and beyond
Liu, S., Deng, N., Sabour, S., Jia, Y., Huang, M., and Mihalcea, R · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Tokenization and the noiseless channel
Zouhar, V., Meister, C., Gastaldi, J., Du, L., Sachan, M., and Cotterell, R · 2023
Cited alongside, same era.
Liu, Y., Lin, P., Wang, M., and Schütze, H
Cited in the paper.
Minixhofer, B., Ponti, E. M., and Vulić, I · 2024
Closest in time.
Language model tokenizers introduce unfairness between languages
Petrov, A., La Malfa, E., Torr, P., and Bibi, A · 2024
Closest in time.
Toward a theory of tokenization in llms
Rajaraman, N., Jiao, J., and Ramchandran, K · 2024
Closest in time.
Tokenization is more than compression
Schmidt, C. W., Reddy, V., Zhang, H., Alameddine, A., Uzan, O., Pinter, Y., and Tanner, C · 2024
Closest in time.
Tokenization counts: the impact of tokenization on arithmetic in frontier llms
Singh, A. K. and Strouse, D · 2024
Closest in time.
Megabyte: Predicting million-byte sequences with multiscale transformers
Yu, L., Simig, D., Flaherty, C., Aghajanyan, A., Zettlemoyer, L., and Lewis, M · 2024
Closest in time.