Fetching the paper…
Reading the bibliography…
Tokenization - the practice of converting strings of characters from an alphabet into sequences of tokens over a vocabulary - is a critical step in the NLP pipeline.
Transductions and Context-Free Languages
Jean Berstel · 1979
Earlier work this paper cites.
A generalization of Ginsburg and Rose’s characterization of G-S-M mappings
C. Choffrut · 1979
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage · 1994
Earlier work this paper cites.
Critical tokenization and its properties
Jin Guo · 1997
Earlier work this paper cites.
Theory of Point Estimation
Erich L. Lehmann and George Casella · 1998
Earlier work this paper cites.
“Maximal-munch” tokenization in linear time
Thomas Reps · 1998
Earlier work this paper cites.
Tokenisation and sentence segmentation
David D. Palmer · 2000
Earlier work this paper cites.
Unsupervised discovery of morphemes
Mathias Creutz and Krista Lagus · 2002
Earlier work this paper cites.
Differentiable weighted finite-state transducers, 2020
Awni Hannun, Vineel Pratap, Jacob Kahn, and Wei-Ning Hsu · 2010
Earlier work this paper cites.
Japanese and Korean voice search
Mike Schuster and Kaisuke Nakajima · 2012
Earlier work this paper cites.
A Bayesian characterization of relative entropy
John C. Baez and Tobias Fritz · 2014
Earlier work this paper cites.
On using very large target vocabulary for neural machine translation
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio · 2015
Earlier work this paper cites.
Addressing the rare word problem in neural machine translation
Thang Luong, Ilya Sutskever, Quoc Le, Oriol Vinyals, and Wojciech Zaremba · 2015
Earlier work this paper cites.
Achieving open vocabulary neural machine translation with hybrid word-character models
Minh-Thang Luong and Christopher D. Manning · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Z. Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason R. Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Gregory S. Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Earlier work this paper cites.
Multiscale sequence modeling with a learned dictionary, 2017
Bart van Merriënboer, Amartya Sanyal, Hugo Larochelle, and Yoshua Bengio · 2017
Earlier work this paper cites.
Neural lattice language models
Jacob Buckman and Graham Neubig · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo · 2018
Cited alongside, same era.
A call for prudent choice of subword merge operations in neural machine translation
Shuoyang Ding, Adithya Renduchintala, and Kevin Duh · 2019
Cited alongside, same era.
Training hybrid language models by marginalizing over segmentations
Edouard Grave, Sainbayar Sukhbaatar, Piotr Bojanowski, and Armand Joulin · 2019
Cited alongside, same era.
Spell once, summon anywhere: A two-level open-vocabulary language model
Sabrina J. Mielke and Jason Eisner · 2019
Cited alongside, same era.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett · 2020
Cited alongside, same era.
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita · 2020
How much does tokenization affect neural machine translation?
Miguel Domingo, Mercedes García-Martínez, Alexandre Helle, Francisco Casacuberta, and Manuel Herranz · 2023
Later among the works it cites.
How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in Japanese
Takuro Fujii, Koki Shibata, Atsuki Yamaguchi, Terufumi Morishita, and Yasuhiro Sogawa · 2023
Later among the works it cites.
Effects of sub-word segmentation on performance of transformer language models
Jue Hou, Anisia Katinskaia, Anh-Duc Vu, and Roman Yangarber · 2023
Later among the works it cites.
A better way to do masked language model scoring
Carina Kauf and Anna Ivanova · 2023
Later among the works it cites.
HuggingFace’s Tokenizers, April 2023
Anthony Moi and Nicolas Patry · 2023
Later among the works it cites.
Efficient transformers with dynamic token pooling
Piotr Nawrot, Jan Chorowski, Adrian Lancucki, and Edoardo Maria Ponti · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Masked language model scoring
Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff · 2020
Cited alongside, same era.
You should evaluate your language model on marginal likelihood over tokenisations
Kris Cao and Laura Rimell · 2021
Cited alongside, same era.
Superbizarre is not superb: Derivational morphology improves BERT’s interpretation of complex words
Valentin Hofmann, Janet Pierrehumbert, and Hinrich Schütze · 2021
Cited alongside, same era.
Between words and characters: A brief history of open-vocabulary modeling and tokenization in NLP
Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y. Lee, Benoît Sagot, and Samson Tan · 2021
Cited alongside, same era.
Fast WordPiece tokenization
Xinying Song, Alex Salcianu, Yang Song, Dave Dopson, and Denny Zhou · 2021
Cited alongside, same era.
Canine: Pre-training an efficient tokenization-free encoder for language representation
Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting · 2022
Cited alongside, same era.
Later among the works it cites.
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell · 2023
Later among the works it cites.
A formal perspective on byte-pair encoding
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Tim Vieira, Mrinmaya Sachan, and Ryan Cotterell · 2023
Later among the works it cites.
Token alignment via character matching for subword completion, 2024
Ben Athiwaratkun, Shiqi Wang, Mingyue Shang, Yuchen Tian, Zijian Wang, Sujan Kumar Gonugondla, Sanjay Krishna Gouda, Rob Kwiatowski, Ramesh Nallapati, and Bing Xiang · 2024
Closest in time.
Constructing a BPE tokenization DFA
Martin Berglund, Willeke Martens, and Brink van der Merwe · 2024
Closest in time.
On the Proper Treatment of Tokenization in Psycholinguistic
Mario Giulianelli, Luca Malagutti, Juan Luis Gastaldi, Brian DuSell, Tim Vieira, and Ryan Cotterell · 2024
Closest in time.
Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition
Daniel Jurafsky and James H. Martin · 2024
Closest in time.
Toward a theory of tokenization in LLMs, 2024
Nived Rajaraman, Jiantao Jiao, and Kannan Ramchandran · 2024
Closest in time.
Tokenization is more than compression, 2024
Craig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner · 2024
Closest in time.
Greed is all you need: An evaluation of tokenizer inference methods, 2024
Omri Uzan, Craig W. Schmidt, Chris Tanner, and Yuval Pinter · 2024
Closest in time.
From Language Models over Tokens to Language Models over Characters, 2024
Tim Vieira, Ben LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Brian DuSell, John Terilla, Timothy J. O’Donnell, and Ryan Cotterell · 2024
Closest in time.
MambaByte: Token-free Selective State Space Model, 2024
Junxiong Wang, Tushaar Gangavarapu, Jing Nathan Yan, and Alexander M. Rush · 2024
Closest in time.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst · 2045
Closest in time.