Fetching the paper…
Reading the bibliography…
In this paper, we introduce ELECTRA-style tasks to cross-lingual language model pre-training.
SpanBERT: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2019 · 1907
Earlier work this paper cites.
WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2019 · 1907
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017 · 1958
Earlier work this paper cites.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020b · 2003
Earlier work this paper cites.
A systematic comparison of various statistical alignment models
Franz Josef Och and Hermann Ney. 2003 · 2003
Earlier work this paper cites.
Explicit alignment objectives for multilingual bidirectional encoders
Junjie Hu, Melvin Johnson, Orhan Firat, Aditya Siddhant, and Graham Neubig. 2020a · 2010
Earlier work this paper cites.
VECO: Variable encoder-decoder pre-training for cross-lingual understanding and generation
Fuli Luo, Wei Wang, Jiahao Liu, Yijia Liu, Bin Bi, Songfang Huang, Fei Huang, and Luo Si. 2020 · 2010
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
A simple, fast, and effective reparameterization of ibm model 2
Chris Dyer, Victor Chahuneau, and Noah A Smith. 2013 · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
A decomposable attention model for natural language inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Earlier work this paper cites.
The united nations parallel corpus v1. 0
Michał Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
The IIT Bombay English-Hindi parallel corpus
Anoop Kunchukuttan, Pratik Mehta, and Pushpak Bhattacharyya. 2018 · 2018
Cited alongside, same era.
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
Mikel Artetxe and Holger Schwenk. 2019 · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Unified language model pre-training for natural language understanding and generation
CCAligned: A massive collection of cross-lingual web-document pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, and Philipp Koehn. 2020 · 2020
Later among the works it cites.
SimAlign: High quality word alignments without parallel training data using static and contextualized embeddings
Masoud Jalili Sabet, Philipp Dufter, François Yvon, and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
MLQA: Evaluating cross-lingual extractive question answering
Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Alternating language modeling for cross-lingual pre-training
Jian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu, Zhoujun Li, and Ming Zhou. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 2019
Cited alongside, same era.
Massively multilingual transfer for NER
Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019 · 2019
Cited alongside, same era.
PAWS-X: A cross-lingual adversarial dataset for paraphrase identification
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019a · 2019
Cited alongside, same era.
Universal dependencies 2.5
Daniel Zeman, Joakim Nivre, Mitchell Abrams, and et al. 2019 · 2019
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Cited alongside, same era.
UniLMv2: Pseudo-masked language models for unified language model pre-training
Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Jianfeng Gao, Songhao Piao, Ming Zhou, and Hsiao-Wuen Hon. 2020 · 2020
Cited alongside, same era.
Multilingual alignment of contextual word representations
Steven Cao, Nikita Kitaev, and Dan Klein. 2020 · 2020
Cited alongside, same era.
InfoXLM: An information-theoretic framework for cross-lingual language model pre-training
Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, and Ming Zhou. 2021b · 2021
Closest in time.
Larger-scale transformers for multilingual masked language modeling
Naman Goyal, Jingfei Du, Myle Ott, Giri Anantharaman, and Alexis Conneau. 2021 · 2021
Closest in time.
Learning to sample replacements for ELECTRA pre-training
Yaru Hao, Li Dong, Hangbo Bao, Ke Xu, and Furu Wei. 2021 · 2021
Closest in time.
Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, Alexandre Muzio, Saksham Singhal, Hany Hassan Awadalla, Xia Song, and Furu Wei. 2021 · 2021
Closest in time.
COCO-LM: Correcting and contrasting text sequences for language model pretraining
Yu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary, Paul Bennett, Jiawei Han, and Xia Song. 2021 · 2021
Closest in time.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Closest in time.
Inducing language-agnostic multilingual representations
Wei Zhao, Steffen Eger, Johannes Bjerva, and Isabelle Augenstein. 2021 · 2021
Closest in time.
Allocating large vocabulary capacity for cross-lingual language model pre-training
Bo Zheng, Li Dong, Shaohan Huang, Saksham Singhal, Wanxiang Che, Ting Liu, Xia Song, and Furu Wei. 2021 · 2021
Closest in time.
Corrupted image modeling for self-supervised visual pre-training
Yuxin Fang, Li Dong, Hangbo Bao, Xinggang Wang, and Furu Wei. 2022 · 2022
Closest in time.