Fetching the paper…
Reading the bibliography…
While subword tokenizers such as BPE and WordPiece are typically used to build vocabularies for NLP models, the method of decoding text into a sequence of tokens from these vocabularies is often left unspecified, or ill-suited to the method in which they were constructed.
Morpheme segmentation gold standards for finnish and english
Mathias Creutz and Bo Krister Johan Linden. 2004 · 2004
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Z. Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason R. Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Gregory S. Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Earlier work this paper cites.
Morpholex: A derivational morphological database for 70,000 english words
Claudia H. Sánchez-Gutiérrez, Hugo Mailhot, S. Hélène Deacon, and Maximiliano A. Wilson. 2018 · 2018
Earlier work this paper cites.
Ladec: The large database of english compounds
Christina L. Gagné, Thomas L. Spalding, and Daniel Schmidtke. 2019 · 2019
Earlier work this paper cites.
Investigating the effectiveness of BPE: The power of shorter sequences
Matthias Gallé. 2019 · 2019
Earlier work this paper cites.
Character eyes: Seeing language through character-level taggers
Yuval Pinter, Marc Marone, and Jacob Eisenstein. 2019 · 2019
Earlier work this paper cites.
Rare words: A major problem for contextualized embeddings and how to fix it by attentive mimicking
Timo Schick and Hinrich Schütze. 2019 · 2019
Earlier work this paper cites.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Earlier work this paper cites.
Finding the optimal vocabulary size for neural machine translation
Thamme Gowda and Jonathan May. 2020 · 2020
Earlier work this paper cites.
Dynamic programming encoding for subword segmentation in neural machine translation
Xuanli He, Gholamreza Haffari, and Mohammad Norouzi. 2020 · 2020
Earlier work this paper cites.
DagoBERT: Generating derivational morphology with a pretrained language model
Valentin Hofmann, Janet Pierrehumbert, and Hinrich Schütze. 2020 · 2020
Earlier work this paper cites.
Getting the ##life out of living: How adequate are word-pieces for modelling complex morphology?
Stav Klein and Reut Tsarfaty. 2020 · 2020
Cited alongside, same era.
Domain adaptation challenges of BERT in tokenization and sub-word representations of out-of-vocabulary words
Anmol Nayak, Hariprasad Timmapathini, Karthikeyan Ponnalagu, and Vijendran Gopalan Venkoparao. 2020 · 2020
Cited alongside, same era.
Will it unblend?
Yuval Pinter, Cassandra L. Jacobs, and Jacob Eisenstein. 2020 · 2020
Cited alongside, same era.
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020 · 2020
Cited alongside, same era.
MorphyNet: a large multilingual database of derivational and inflectional morphology
Khuyagbaatar Batsuren, Gábor Bella, and Fausto Giunchiglia. 2021 · 2021
Cited alongside, same era.
From characters to words: the turning point of BPE merges
BPE-knockout: Systematic review of BPE tokenisers and their flaws with application in Dutch morphology
Thomas Bauwens. 2023 · 2023
Later among the works it cites.
Analyzing cognitive plausibility of subword tokenization
Lisa Beinborn and Yuval Pinter. 2023 · 2023
Later among the works it cites.
The minipile challenge for data-efficient language models
Jean Kaddour. 2023 · 2023
Later among the works it cites.
XLM-V: Overcoming the vocabulary bottleneck in multilingual masked language models
Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, and Madian Khabsa. 2023 · 2023
Later among the works it cites.
CompoundPiece: Evaluating and improving decompounding performance of language models
Benjamin Minixhofer, Jonas Pfeiffer, and Ivan Vulić. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ximena Gutierrez-Vasques, Christian Bentz, Olga Sozinova, and Tanja Samardzic. 2021 · 2021
Cited alongside, same era.
Superbizarre is not superb: Derivational morphology improves BERT’s interpretation of complex words
Valentin Hofmann, Janet Pierrehumbert, and Hinrich Schütze. 2021 · 2021
Cited alongside, same era.
Between words and characters: a brief history of open-vocabulary modeling and tokenization in nlp
Sabrina J Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y Lee, Benoît Sagot, et al. 2021 · 2021
Cited alongside, same era.
UniMorph 4.0: Universal Morphology
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieraś, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Abbott Lane, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóǧa, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan North Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Prud’hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya D. McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, and Ekaterina Vylomova. 2022 · 2022
Cited alongside, same era.
Improving tokenisation by alternative treatment of spaces
Edward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, and Aline Villavicencio. 2022 · 2022
Cited alongside, same era.
An embarrassingly simple method to mitigate undesirable properties of pretrained language model tokenizers
Valentin Hofmann, Hinrich Schuetze, and Janet Pierrehumbert. 2022 · 2022
Cited alongside, same era.
Cassandra L Jacobs and Yuval Pinter. 2022 · 2022
Cited alongside, same era.
What changes when you randomly choose BPE merge operations? not much
Jonne Saleva and Constantine Lignos. 2023 · 2023
Later among the works it cites.
Incorporating context into subword vocabularies
Shaked Yehezkel and Yuval Pinter. 2023 · 2023
Later among the works it cites.
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
Evaluating subword tokenization: Alien subword composition and oov generalization challenge
Khuyagbaatar Batsuren, Ekaterina Vylomova, Verna Dankers, Tsetsuukhei Delgerbaatar, Omri Uzan, Yuval Pinter, and Gábor Bella. 2024 · 2024
Closest in time.
Two counterexamples to tokenization and the noiseless channel
Marco Cognetta, Vilém Zouhar, Sangwhan Moon, and Naoaki Okazaki. 2024b · 2024
Closest in time.
Word boundary information isn’t useful for encoder language models
Edward Gow-Smith, Dylan Phelps, Harish Tayyar Madabushi, Carolina Scarton, and Aline Villavicencio. 2024 · 2024
Closest in time.
Tokenization matters: Navigating data-scarce tokenization for gender inclusive language technologies
Anaelia Ovalle, Ninareh Mehrabi, Palash Goyal, Jwala Dhamala, Kai-Wei Chang, Richard Zemel, Aram Galstyan, Yuval Pinter, and Rahul Gupta. 2024 · 2024
Closest in time.
Tokenization is more than compression
Craig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner. 2024 · 2024
Closest in time.