Fetching the paper…
Reading the bibliography…
Tokenization is a fundamental preprocessing step for almost all NLP tasks.
Trie memory
Edward Fredkin. 1960 · 1960
Earlier work this paper cites.
Error bounds for convolutional codes and an asymptotically optimum decoding algorithm
A. Viterbi. 1967 · 1967
Earlier work this paper cites.
Efficient string matching: An aid to bibliographic search
Alfred V. Aho and Margaret J. Corasick. 1975 · 1975
Earlier work this paper cites.
Basic engineering for Chinese processing–modern Chinese word frequency count
Yuan Liu and Nanyuan Liang. 1986 · 1986
Earlier work this paper cites.
On the Methods of Chinese Automatic Segmentation
Chunyu Jie, Yuan Liu, and Nanyuan Liang. 1989 · 1989
Earlier work this paper cites.
Tokenization as the initial phase in NLP
Jonathan J. Webster and Chunyu Kit. 1992 · 1992
Earlier work this paper cites.
Regular models of phonological rule systems
Ronald M. Kaplan and Martin Kay. 1994 · 1994
Earlier work this paper cites.
Finite-state transducers in language and speech processing
Mehryar Mohri. 1997 · 1997
Earlier work this paper cites.
“Maximal-Munch” Tokenization in Linear Time
Thomas Reps. 1998 · 1998
Cited alongside, same era.
Tokenisation and sentence segmentation
David D. Palmer. 2000 · 2000
Cited alongside, same era.
A compact static double-array keeping character codes
Susumu Yata, Masaki Oono, Kazuhiro Morita, Masao Fuketa, Toru Sumitomo, and Jun ichi Aoe. 2007 · 2007
Cited alongside, same era.
Optimizing Chinese word segmentation for machine translation performance
Pi-Chuan Chang, Michel Galley, and Christopher D. Manning. 2008 · 2008
Cited alongside, same era.
Speech and Language Processing (2nd Edition)
Daniel Jurafsky and James H. Martin. 2009 · 2009
Cited alongside, same era.
Japanese and Korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, and et al. 2020 · 2020
Closest in time.
The WordPiece Algorithm in Open Source BERT
Google. 2018 · 2020
Closest in time.
TensorFlow Text
Google. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deterministic word segmentation using maximum matching with fully lexicalized rules
Manabu Sassano. 2014 · 2014
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Closest in time.
Tokenizers
HuggingFace. 2020 · 2020
Closest in time.