Fetching the paper…
Reading the bibliography…
Multilingual pretrained representations generally rely on subword segmentation algorithms to create a shared multilingual vocabulary.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017 · 1958
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2015 · 2015
Earlier work this paper cites.
Exploring the limits of language modeling
Rafal Józefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Earlier work this paper cites.
Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Fully character-level neural machine translation without explicit segmentation
Jason Lee, Kyunghyun Cho, and Thomas Hofmann. 2017 · 2017
Earlier work this paper cites.
Cross-lingual distillation for text classification
Ruochen Xu and Yiming Yang. 2017 · 2017
Earlier work this paper cites.
Compositional representation of morphologically-rich input for neural machine translation
Duygu Ataman and Marcello Federico. 2018 · 2018
Earlier work this paper cites.
Revisiting character-based neural machine translation with capacity and compression
Colin Cherry, George Foster, Ankur Bapna, Orhan Firat, and Wolfgang Macherey. 2018 · 2018
Earlier work this paper cites.
Semi-supervised sequence modeling with cross-view training
Kevin Clark, Minh-Thang Luong, Christopher D. Manning, and Quoc Le. 2018 · 2018
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc Le. 2018 · 2018
Cited alongside, same era.
Exploring bert’s vocabulary
Judit Ács. 2019 · 2019
Cited alongside, same era.
Crosslingual language model pretraining
Alexis Conneau and Guillaume Lample. 2019 · 2019
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Later among the works it cites.
Parsing with multilingual BERT, a small corpus, and a small treebank
Ethan C. Chau, Lucy H. Lin, and Noah A. Smith. 2020 · 2020
Later among the works it cites.
Improving multilingual models with language-clustered vocabularies
Hyung Won Chung, Dan Garrette, Kiat Chuan Tan, and Jason Riesa. 2020 · 2020
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Later among the works it cites.
Dynamic programming encoding for subword segmentation in neural machine translation
Xuanli He, Gholamreza Haffari, and Mohammad Norouzi. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks
Haoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, and Ming Zhoun. 2019 · 2019
Cited alongside, same era.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Cited alongside, same era.
Multilingual neural machine translation with soft decoupled encoding
Xinyi Wang, Hieu Pham, Philip Arthur, and Graham Neubig. 2019 · 2019
Cited alongside, same era.
Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT
Shijie Wu and Mark Dredze. 2019 · 2019
Cited alongside, same era.
PAWS-X: A cross-lingual adversarial dataset for paraphrase identification
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019 · 2019
Cited alongside, same era.
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Later among the works it cites.
MLQA: Evaluating Cross-lingual Extractive Question Answering
Patrick Lewis, Barlas Oğuz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020 · 2020
Later among the works it cites.
Charbert: Character-aware pre-trained language model
Wentao Ma, Yiming Cui, Chenglei Si, Ting Liu, Shijin Wang, and Guoping Hu. 2020 · 2020
Later among the works it cites.
MAD-X: An Adapter-based Framework for Multi-task Cross-lingual Transfer
Jonas Pfeiffer, Ivan Vuli, Iryna Gurevych, and Sebastian Ruder. 2020 · 2020
Later among the works it cites.
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020 · 2020
Later among the works it cites.
Revisit knowledge distillation: a teacher-free framework
Li Yuan, Francis E.H.Tay, Guilin Li, Tao Wang, and Jiashi Feng. 2020 · 2020
Later among the works it cites.
Ambert: A pre-trained language model with multi-grained tokenization
Xinsong Zhang and Hang Li. 2020 · 2020
Later among the works it cites.