Training hybrid language models by marginalizing over segmentations
Edouard Grave, Sainbayar Sukhbaatar, Piotr Bojanowski, and Armand Joulin. 2019 · 2019
Later among the works it cites.
Stochastic tokenization with a language model for neural text classification
Tatsuya Hiraoka, Hiroyuki Shindo, and Yuji Matsumoto. 2019 · 2019
Later among the works it cites.
Learning to discover, ground and use words with segmental neural language models
Kazuya Kawakami, Chris Dyer, and Phil Blunsom. 2019 · 2019
Later among the works it cites.
Glyce: Glyph-vectors for chinese character representations
Yuxian Meng, Wei Wu, Fei Wang, Xiaoya Li, Ping Nie, Fan Yin, Muyu Li, Qinghong Han, Xiaofei Sun, and Jiwei Li. 2019 · 2019
Later among the works it cites.
What kind of language is hard to language-model?
Sabrina J. Mielke, Ryan Cotterell, Kyle Gorman, Brian Roark, and Jason Eisner. 2019 · 2019
Later among the works it cites.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Squared english word: A method of generating glyph to use super characters for sentiment analysis
Baohua Sun, Lin Yang, Catherine Chi, Wenhan Zhang, and Michael Lin. 2019 · 2019
Later among the works it cites.
Neural machine translation with byte-level subwords
Original
Changhan Wang, Kyunghyun Cho, and Jiatao Gu. 2019 · 2019
Later among the works it cites.
Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT
Shijie Wu and Mark Dredze. 2019 · 2019
Later among the works it cites.
A latent morphology model for open-vocabulary neural machine translation
Duygu Ataman, Wilker Aziz, and Alexandra Birch. 2020 · 2020
Later among the works it cites.
Linguist vs. machine: Rapid development of finite-state morphological grammars
Sarah Beemer, Zak Boston, April Bukoski, Daniel Chen, Princess Dickens, Andrew Gerlach, Torin Hopkins, Parth Anand Jawale, Chris Koski, Akanksha Malhotra, Piyush Mishra, Saliha Muradoglu, Lan Sang, Tyler Short, Sagarika Shreevastava, Elizabeth Spaulding, Testumichi Umada, Beilei Xiang, Changbing Yang, and Mans Hulden. 2020 · 2020
Later among the works it cites.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Later among the works it cites.
Parsing with multilingual BERT, a small corpus, and a small treebank
Ethan C. Chau, Lucy H. Lin, and Noah A. Smith. 2020 · 2020
Later among the works it cites.
Improving multilingual models with language-clustered vocabularies
Hyung Won Chung, Dan Garrette, Kiat Chuan Tan, and Jason Riesa. 2020 · 2020
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Later among the works it cites.
CharacterBERT: Reconciling ELMo and BERT for word-level open-vocabulary representations from characters
Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, Hiroshi Noji, Pierre Zweigenbaum, and Jun’ichi Tsujii. 2020 · 2020
Later among the works it cites.
Finding the optimal vocabulary size for neural machine translation
Thamme Gowda and Jonathan May. 2020 · 2020
Later among the works it cites.
Morfessor EM+Prune: Improved subword segmentation with expectation maximization and pruning
Stig-Arne Grönroos, Sami Virpioja, and Mikko Kurimo. 2020 · 2020
Later among the works it cites.
Dynamic programming encoding for subword segmentation in neural machine translation
Xuanli He, Gholamreza Haffari, and Mohammad Norouzi. 2020 · 2020
Later among the works it cites.
The unstoppable rise of computational linguistics in deep learning
James Henderson. 2020 · 2020
Later among the works it cites.
Optimizing word segmentation for downstream task
Tatsuya Hiraoka, Sho Takase, Kei Uchiumi, Atsushi Keyaki, and Naoaki Okazaki. 2020 · 2020
Later among the works it cites.
Cross-lingual ability of multilingual bert: An empirical study
Karthikeyan K, Zihan Wang, Stephen Mayhew, and Dan Roth. 2020 · 2020
Later among the works it cites.
Towards reasonably-sized character-level transformer NMT by finetuning subword systems
Jindřich Libovický and Alexander Fraser. 2020 · 2020
Later among the works it cites.
CharBERT: Character-aware Pre-trained Language Model
Wentao Ma, Yiming Cui, Chenglei Si, Ting Liu, Shijin Wang, and Guoping Hu. 2020 · 2020
Later among the works it cites.
Towards end-to-end in-image neural machine translation
Elman Mansimov, Mitchell Stern, Mia Chen, Orhan Firat, Jakob Uszkoreit, and Puneet Jain. 2020 · 2020
Later among the works it cites.
An empirical study of tokenization strategies for various korean nlp tasks
Kyubyong Park, Joohong Lee, Seongbo Jang, and Dawoon Jung. 2020 · 2020
Later among the works it cites.
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020 · 2020
Later among the works it cites.
Neural polysynthetic language modelling
Original
Lane Schwartz, Francis Tyers, Lori Levin, Christo Kirov, Patrick Littell, Chi kiu Lo, Emily Prud’hommeaux, Hyunji Hayley Park, Kenneth Steimel, Rebecca Knowles, Jeffrey Micher, Lonny Strunk, Han Liu, Coleman Haley, Katherine J. Zhang, Robbie Jimmerson, Vasilisa Andriyanets, Aldrian Obaja Muis, Naoki Otani, Jong Hyuk Park, and Zhisong Zhang. 2020 · 2020
Later among the works it cites.
Mind your inflections! Improving NLP for non-standard Englishes with Base-Inflection Encoding
Samson Tan, Shafiq Joty, Lav Varshney, and Min-Yen Kan. 2020 · 2020
Later among the works it cites.
Extending multilingual BERT to low-resource languages
Zihan Wang, Karthikeyan K, Stephen Mayhew, and Dan Roth. 2020b · 2020
Later among the works it cites.
Char2subword: Extending the subword embedding space using robust character compositionality
Original
Gustavo Aguilar, Bryan McCann, Tong Niu, Nazneen Rajani, Nitish Keskar, and Thamar Solorio. 2021 · 2021
Closest in time.
Evaluating various tokenizers for arabic text classification
Original
Zaid Alyafeai, Maged S. Al-shaibani, Mustafa Ghaleb, and Irfan Ahmad. 2021 · 2021
Closest in time.
How suitable are subword segmentation strategies for translating non-concatenative morphology?
Chantal Amrhein and Rico Sennrich. 2021 · 2021
Closest in time.
You should evaluate your language model on marginal likelihood over tokenisations
Kris Cao and Laura Rimell. 2021 · 2021
Closest in time.
Canine: Pre-training an efficient tokenization-free encoder for language representation
Original
Jonathan H Clark, Dan Garrette, Iulia Turc, and John Wieting. 2021 · 2021
Closest in time.
Crowdsourced phrase-based tokenization for low-resourced neural machine translation: The case of fon language
Original
Bonaventure F. P. Dossou and Chris C. Emezue. 2021 · 2021
Closest in time.
A masked segmental language model for unsupervised natural language segmentation
Original
C. M. Downey, Fei Xia, Gina-Anne Levow, and Shane Steinert-Threlkeld. 2021 · 2021
Closest in time.
How to adapt your pretrained multilingual model to 1600 languages
Original
Abteen Ebrahimi and Katharina Kann. 2021 · 2021
Closest in time.
How to split: the effect of word segmentation on gender bias in speech translation
Marco Gaido, Beatrice Savoldi, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2021 · 2021
Closest in time.
From characters to words: the turning point of BPE merges
Ximena Gutierrez-Vasques, Christian Bentz, Olga Sozinova, and Tanja Samardzic. 2021 · 2021
Closest in time.
Joint optimization of tokenization and downstream model
Tatsuya Hiraoka, Sho Takase, Kei Uchiumi, Atsushi Keyaki, and Naoaki Okazaki. 2021 · 2021
Closest in time.
Superbizarre is not superb: Derivational morphology improves BERT’s interpretation of complex words
Valentin Hofmann, Janet Pierrehumbert, and Hinrich Schütze. 2021 · 2021
Closest in time.
Models in a spelling bee: Language models implicitly learn the character composition of tokens
Original
Itay Itzhak and Omer Levy. 2021 · 2021
Closest in time.
How bpe affects memorization in transformers
Original
Eugene Kharitonov, Marco Baroni, and Dieuwke Hupkes. 2021 · 2021
Closest in time.
Why don’t people use character-level machine translation?
Original
Jindřich Libovický, Helmut Schmid, and Alexander Fraser. 2021 · 2021
Closest in time.
Bridging subword gaps in pretrain-finetune paradigm for natural language generation
Original
Xin Liu, Baosong Yang, Dayiheng Liu, Haibo Zhang, Weihua Luo, Min Zhang, Haiying Zhang, and Jinsong Su. 2021 · 2021
Closest in time.
Recurrent neural networks with mixed hierarchical structures for natural language processing
Original
Zhaoxin Luo and Michael Zhu. 2021 · 2021
Closest in time.
Wine is not v i n. on the compatibility of tokenizations across languages
Antonis Maronikolakis, Philipp Dufter, and Hinrich Schütze. 2021 · 2021
Closest in time.
One size does not fit all: Finding the optimal n-gram sizes for fasttext models across languages
Original
Vít Novotný, Eniafe Festus Ayetiran, Dávid Lupták, Michal Stefánik, and Petr Sojka. 2021 · 2021
Closest in time.
Integrating approaches to word representation
Original
Yuval Pinter. 2021 · 2021
Closest in time.
Noisy UGC translation at the character level: Revisiting open-vocabulary capabilities and robustness of char-based models
José Carlos Rosales Núñez, Guillaume Wisniewski, and Djamé Seddah. 2021 · 2021
Closest in time.
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021 · 2021
Closest in time.
Robust open-vocabulary translation from visual text representations
Elizabeth Salesky, David Etter, and Matt Post. 2021 · 2021
Closest in time.
The effectiveness of morphology-aware segmentation in low-resource neural machine translation
Jonne Saleva and Constantine Lignos. 2021 · 2021
Closest in time.
ShuoWen-JieZi: Linguistically Informed Tokenizers For Chinese Language Model Pretraining
Original
Chenglei Si, Zhengyan Zhang, Yingfa Chen, Fanchao Qi, Xiaozhi Wang, Zhiyuan Liu, and Maosong Sun. 2021 · 2021
Closest in time.
Charformer: Fast character transformers via gradient-based subword tokenization
Original
Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler. 2021 · 2021
Closest in time.
A statistical extension of byte-pair encoding
David Vilar and Marcello Federico. 2021 · 2021
Closest in time.
Multi-view subword regularization
Xinyi Wang, Sebastian Ruder, and Graham Neubig. 2021b · 2021
Closest in time.
Vocabulary learning via optimal transport for neural machine translation
Jingjing Xu, Hao Zhou, Chun Gan, Zaixiang Zheng, and Lei Li. 2021 · 2021
Closest in time.
Byt5: Towards a token-free future with pre-trained byte-to-byte models
Original
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2021 · 2021
Closest in time.
AMBERT: A pre-trained language model with multi-grained tokenization
Xinsong Zhang, Pengshuai Li, and Hang Li. 2021 · 2021
Closest in time.
Learning character-level compositionality with visual features
Frederick Liu, Han Lu, Chieh Lo, and Graham Neubig. 2017 · 2068
Closest in time.
Fast WordPiece tokenization
Xinying Song, Alex Salcianu, Yang Song, Dave Dopson, and Denny Zhou. 2021 · 2089
Closest in time.