Fetching the paper…
Reading the bibliography…
Tokenization is an important preprocessing step in the training and inference of large language models (LLMs).
Introduction to Automata Theory, Languages, and Computation (3rd Edition)
John E. Hopcroft, Rajeev Motwani, and Jeffrey D. Ullman. 2006 · 2006
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates
Taku Kudo. 2018 · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
A general-purpose algorithm for constrained sequential inference
Daniel Deutsch, Shyam Upadhyay, and Dan Roth. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Neural machine translation with byte-level subwords
Changhan Wang, Kyunghyun Cho, and Jiatao Gu. 2019 · 2019
Cited alongside, same era.
Byte Pair Encoding is Suboptimal for Language Model Pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Cited alongside, same era.
Constrained language models yield few-shot semantic parsers
Richard Shin, Christopher Lin, Sam Thomson, Charles Chen, Subhro Roy, Emmanouil Antonios Platanios, Adam Pauls, Dan Klein, Jason Eisner, and Benjamin Van Durme. 2021 · 2021
Cited alongside, same era.
Validating large language models with relm
Michael Kuchnik, Virginia Smith, and George Amvrosiadis. 2023 · 2023
Cited alongside, same era.
Grammar prompting for domain-specific language generation with large language models
Bailin Wang, Zi Wang, Xuezhi Wang, Yuan Cao, Rif A. Saurous, and Yoon Kim. 2023 · 2023
Cited alongside, same era.
Efficient guided generation for large language models
Brandon T. Willard and Rémi Louf. 2023 · 2023
Later among the works it cites.
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
Grammar-constrained decoding for structured nlp tasks without finetuning
Saibo Geng, Martin Josifoski, Maxime Peyrard, and Robert West. 2024 · 2024
Closest in time.
Guidance
guidance-ai. 2024 · 2024
Closest in time.
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
Aaditya K. Singh and D. J. Strouse. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…