Fetching the paper…
Reading the bibliography…
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
MASS: masked sequence to sequence pre-training for language generation
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019 · 1905
Earlier work this paper cites.
Revisiting self-training for neural sequence generation
Junxian He, Jiatao Gu, Jiajun Shen, and Marc’Aurelio Ranzato. 2019 · 1909
Earlier work this paper cites.
The source-target domain mismatch problem in machine translation
Jiajun Shen, Peng-Jen Chen, Matt Le, Junxian He, Jiatao Gu, Myle Ott, Michael Auli, and Marc’Aurelio Ranzato. 2019 · 1909
Earlier work this paper cites.
Facebook ai’s WAT19 myanmar-english translation task submission
Peng-Jen Chen, Jiajun Shen, Matt Le, Vishrav Chaudhary, Ahmed El-Kishky, Guillaume Wenzek, Myle Ott, and Marc’Aurelio Ranzato. 2019 · 1910
Earlier work this paper cites.
Transformers without tears: Improving the normalization of self-attention
Toan Q. Nguyen and Julian Salazar. 2019a · 1910
Earlier work this paper cites.
Bpe-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2019 · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 1910
Earlier work this paper cites.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Edouard Grave, Michael Auli, and Armand Joulin. 2020 · 2010
Earlier work this paper cites.
Wikilingua: A new benchmark dataset for cross-lingual abstractive summarization
Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. 2020 · 2010
Earlier work this paper cites.
Iterative, mt-based sentence alignment of parallel texts
Rico Sennrich and Martin Volk. 2011 · 2011
Earlier work this paper cites.
Parallel data, tools and interfaces in opus
Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
The AMARA corpus: Building parallel language resources for the educational domain
Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, and Stephan Vogel. 2014 · 2014
Earlier work this paper cites.
Building subject-aligned comparable corpora and mining it for truly parallel sentence pairs
Krzysztof Wołk and Krzysztof Marasek. 2014 · 2014
Cited alongside, same era.
The IWSLT 2015 Evaluation Campaign
M. Cettolo, J. Niehues, S. Stüker, L. Bentivogli, R. Cattoni, and Marcello Federico. 2015a · 2015
Cited alongside, same era.
The IWSLT 2015 evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, Roldano Cattoni, and Marcello Federico. 2015b · 2015
Cited alongside, same era.
A massively parallel corpus: the bible in 100 languages
Christos Christodouloupoulos and Mark Steedman. 2015 · 2015
Cited alongside, same era.
OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles
Pierre Lison and Jörg Tiedemann. 2016 · 2016
Cited alongside, same era.
Introducing the Asian language treebank (ALT)
Ye Kyaw Thu, Win Pa Pa, Masao Utiyama, Andrew Finch, and Eiichiro Sumita. 2016 · 2016
ParaCrawl: Web-scale acquisition of parallel corpora
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, and Jaume Zaragoza. 2020 · 2020
Later among the works it cites.
Vietnam’s Development Success Story and the Unfinished SDG Agenda
Anja Baum. 2020 · 2020
Later among the works it cites.
CCAligned: A massive collection of cross-lingual web-document pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, and Philipp Koehn. 2020 · 2020
Later among the works it cites.
Goals, Challenges and Findings of the VLSP 2020 English-Vietnamese News Translation Shared Task
Thanh-Le Ha, Van-Khanh Tran, and Kim-Anh Nguyen. 2020 · 2020
Later among the works it cites.
Multilingual Denoising Pre-training for Neural Machine Translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Semi-supervised sequence modeling with cross-view training
Kevin Clark, Minh-Thang Luong, Christopher D. Manning, and Quoc V. Le. 2018 · 2018
Cited alongside, same era.
Understanding back-translation at scale
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018 · 2018
Cited alongside, same era.
Universal neural machine translation for extremely low resource languages
Jiatao Gu, Hany Hassan, Jacob Devlin, and Victor O.K. Li. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Prompsit’s submission to wmt 2018 parallel corpus filtering shared task
Víctor M. Sánchez-Cartagena, Marta Bañón, Sergio Ortiz-Rojas, and Gema Ramírez-Sánchez · 2018
Cited alongside, same era.
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020 · 2020
Later among the works it cites.
Making monolingual sentence embeddings multilingual using knowledge distillation
Nils Reimers and Iryna Gurevych. 2020 · 2020
Later among the works it cites.
CCNet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020 · 2020
Later among the works it cites.
Improving large-scale language models and resources for filipino
Jan Christian Blaise Cruz and Charibeth Cheng. 2021 · 2021
Later among the works it cites.
PhoMT: A high-quality and large-scale benchmark dataset for Vietnamese-English machine translation
Long Doan, Linh The Nguyen, Nguyen Luong Tran, Thai Hoang, and Dat Quoc Nguyen. 2021a · 2021
Later among the works it cites.
Styled augmented translation (sat)
Chinh Ngo and Trieu H. Trinh. 2021 · 2021
Later among the works it cites.
Geographical distance is the new hyperparameter: A case study of finding the optimal pre-trained language for english-isizulu machine translation
Muhammad Umair Nasir and Innocent Amos Mchechesi. 2022 · 2022
Closest in time.
Vit5: Pretrained text-to-text transformer for vietnamese language generation
Long Phan, Hieu Tran, Hieu Nguyen, and Trieu H. Trinh. 2022 · 2022
Closest in time.