Fetching the paper…
Reading the bibliography…
In this paper, we formalize practical byte pair encoding tokenization as it is used in large language models and other NLP systems, in particular we formally define and investigate the semantics of the SentencePiece and HuggingFace tokenizers, in particular how they relate to each other, depending on how the tokenization rules are constructed.
Computational linguistics
Marti A Hearst (1997): Text tiling: Segmenting text into multi-paragraph subtopic passages · 1997
Earlier work this paper cites.
Pearson Education
Monica Lam, Ravi Sethi, Jeffrey D Ullman & Alfred Aho (2006): Compilers: principles, techniques, and tools · 2006
Earlier work this paper cites.
In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Rico Sennrich, Barry Haddow & Alexandra Birch (2016): Neural Machine Translation of Rare Words with Subword Units · 2016
Earlier work this paper cites.
In Yo-Sub Han & Kai Salomaa, editors: Implementation and Application of Automata - 21st International Conference, CIAA 2016, Seoul, South Korea, July 19-22, 2016, Proceedings
Nicolaas Weideman, Brink van der Merwe, Martin Berglund & Bruce W. Watson (2016): Analyzing Matching Time Behavior of Backtracking Regular Expression Matchers by Using Ambiguity of NFA · 2016
Cited alongside, same era.
In: The World Wide Web Conference
Philippe Skolka, Cristian-Alexandru Staicu & Michael Pradel (2019): Anything to Hide? Studying Minified and Obfuscated Code in the Web · 2019
Cited alongside, same era.
Available at https://openai.com/blog/chatgpt/
OpenAI (2022): ChatGPT: Optimizing language models for dialogue · 2022
Cited alongside, same era.
arXiv: https://arxiv.org/abs/2305.12987
Ariel Ekgren, Amaru Cuba Gyllensten, Felix Stollenwerk, Joey Öhman, Tim Isbister, Evangelia Gogoulou, Fredrik Carlsson, Alice Heiman, Judit Casademont & Magnus Sahlgren (2023): GPT-SW3: An Autoregressive Language Model for the Nordic Languages · 2023
Closest in time.
https://github.com/huggingface/transformers/blob/v4.28.1/src/transformers/models/gpt2/tokenization_gpt2.py
Hugging Face (2023): Transformers · 2023
Closest in time.
arXiv: https://arxiv.org/abs/2304.14780
Felix Stollenwerk (2023): Training and Evaluation of a Multilingual Tokenizer for GPT-SW3 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…