Fetching the paper…
Reading the bibliography…
Modern language models are internally -- and mathematically -- distributions over $\it{token}$ strings rather than $\it{character}$ strings, posing numerous challenges for programmers building user applications on top of them.
A mathematical theory of communication
Shannon, C. E · 1948
Earlier work this paper cites.
A new algorithm for data compression
Gage, P · 1994
Earlier work this paper cites.
A probabilistic Earley parser as a psycholinguistic model
Hale, J · 2001
Earlier work this paper cites.
Expectation-based syntactic comprehension
Levy, R · 2008
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Černocký, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Kudo, T · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
BPE-Dropout: Simple and effective subword regularization
Provilkov, I., Emelianenko, D., and Voita, E · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A · 2020
Cited alongside, same era.
You should evaluate your language model on marginal likelihood over tokenisations
Cao, K. and Rimell, L · 2021
Cited alongside, same era.
PICARD: parsing incrementally for constrained auto-regressive decoding from language models
Scholak, T., Schucher, N., and Bahdanau, D · 2021
Cited alongside, same era.
Synchromesh: Reliable code generation from pre-trained language models
Poesia, G., Polozov, A., Le, V., Tiwari, A., Soares, G., Meek, C., and Gulwani, S · 2022
Cited alongside, same era.
Guidance
Microsoft · 2023
Later among the works it cites.
Efficient guided generation for large language models
Willard, B. T. and Louf, R · 2023
Later among the works it cites.
On the proper treatment of tokenization in psycholinguistics
Giulianelli, M., Malagutti, L., Gastaldi, J. L., DuSell, B., Vieira, T., and Cotterell, R · 2024
Closest in time.
Automata-based constraints for language model decoding
Koo, T., Liu, F., and He, L · 2024
Closest in time.
Oh, B.-D. and Schuler, W · 2024
Closest in time.
Understanding and mitigating tokenization bias in language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Should you marginalize over possible tokenizations?
Chirkova, N., Kruszewski, G., Rozen, J., and Dymetman, M · 2023
Cited alongside, same era.
Grammar-constrained decoding for structured NLP tasks without finetuning
Geng, S., Josifoski, M., Peyrard, M., and West, R · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with PagedAttention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
The art of prompt design: Prompt boundaries and token healing
Lundberg, S. and Ribeiro, M. T · 2023
Cited alongside, same era.
Phan, B., Havasi, M., Muckley, M. J., and Ullrich, K · 2024
Closest in time.
How to compute the probability of a word
Pimentel, T. and Meister, C · 2024
Closest in time.
The foundations of tokenization: Statistical and computational concerns
Gastaldi, J. L., Terilla, J., Malagutti, L., DuSell, B., Vieira, T., and Cotterell, R · 2025
Closest in time.
Syntactic and semantic control of large language models via sequential Monte Carlo
Loula, J., LeBrun, B., Du, L., Lipkin, B., Pasti, C., Grand, G., Liu, T., Emara, Y., Freedman, M., Eisner, J., Cotterell, R., Mansinghka, V., Lew, A. K., Vieira, T., and O’Donnell, T. J · 2025
Closest in time.