Fetching the paper…
Reading the bibliography…
We explore threshold vocabulary trimming in Byte-Pair Encoding subword tokenization, a postprocessing step that replaces rare subwords with their component subwords.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn. 2005 · 2005
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico. 2014 · 2014
Earlier work this paper cites.
Fast and robust neural network joint models for statistical machine translation
Jacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard Schwartz, and John Makhoul. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
Montreal neural machine translation systems for WMT’15
Sébastien Jean, Orhan Firat, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
subword-nmt
Rico Sennrich. 2015 · 2015
Earlier work this paper cites.
Efficient softmax approximation for gpus
Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Massive exploration of neural machine translation architectures
Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc Le. 2017 · 2017
Cited alongside, same era.
The University of Edinburgh’s neural MT systems for WMT17
Rico Sennrich, Alexandra Birch, Anna Currey, Ulrich Germann, Barry Haddow, Kenneth Heafield, Antonio Valerio Miceli Barone, and Philip Williams. 2017 · 2017
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Cited alongside, same era.
Investigating the effectiveness of BPE: The power of shorter sequences
Matthias Gallé. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020 · 2020
Later among the works it cites.
Between words and characters: A brief history of open-vocabulary modeling and tokenization in NLP
Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y. Lee, Benoît Sagot, and Samson Tan. 2021 · 2021
Later among the works it cites.
Canine: Pre-training an efficient tokenization-free encoder for language representation
Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. 2022 · 2022
Later among the works it cites.
Improving tokenisation by alternative treatment of spaces
Edward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, and Aline Villavicencio. 2022 · 2022
Later among the works it cites.
An embarrassingly simple method to mitigate undesirable properties of pretrained language model tokenizers
Valentin Hofmann, Hinrich Schuetze, and Janet Pierrehumbert. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Revisiting low-resource neural machine translation: A case study
Rico Sennrich and Biao Zhang. 2019 · 2019
Cited alongside, same era.
Finding the optimal vocabulary size for neural machine translation
Thamme Gowda and Jonathan May. 2020 · 2020
Cited alongside, same era.
Getting the ##life out of living: How adequate are word-pieces for modelling complex morphology?
Stav Klein and Reut Tsarfaty. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
ByT5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022 · 2022
Later among the works it cites.
Analyzing cognitive plausibility of subword tokenization
Lisa Beinborn and Yuval Pinter. 2023 · 2023
Later among the works it cites.
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
Greed is all you need: An evaluation of tokenizer inference methods
Omri Uzan, Craig W. Schmidt, Chris Tanner, and Yuval Pinter. 2024 · 2024
Closest in time.