Fetching the paper…
Reading the bibliography…
State-of-the-art machine translation (MT) systems are typically trained to generate the "standard" target language; however, many languages have multiple varieties (regional varieties, dialects, sociolects, non-native varieties) that are different from the standard language.
A Theory of Justice
John Rawls. 1999 · 1999
Earlier work this paper cites.
A machine translation system between a pair of closely related languages
Kemal Altintas and Ilyas Cicekli. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
The United Nations parallel corpus v1.0
Michał Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016 · 2002
Earlier work this paper cites.
Can we translate letters?
David Vilar, Jan-T. Peter, and Hermann Ney. 2007 · 2007
Earlier work this paper cites.
Design of the Moses decoder for statistical machine translation
Hieu Hoang and Philipp Koehn. 2008 · 2008
Earlier work this paper cites.
Character-based PSMT for closely related languages
Jörg Tiedemann. 2009 · 2009
Earlier work this paper cites.
BP2EP - adaptation of Brazilian Portuguese texts to European Portuguese
Luis Marujo, Nuno Grazina, Tiago Luis, Wang Ling, Luisa Coheur, and Isabel Trancoso. 2011 · 2011
Earlier work this paper cites.
WIT3: Web inventory of transcribed and translated talks
Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012 · 2012
Earlier work this paper cites.
Combining word-level and character-level models for machine translation between closely-related languages
Preslav Nakov and Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
A malay dialect translation and synthesis system: Proposal and preliminary system
T. Tan, S. Goh, and Y. Khaw. 2012 · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Copied monolingual data improves low-resource neural machine translation
Anna Currey, Antonio Valerio Miceli Barone, and Kenneth Heafield. 2017 · 2017
Earlier work this paper cites.
OpenNMT: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander Rush. 2017 · 2017
Earlier work this paper cites.
Deciphering related languages
Nima Pourdamghani and Kevin Knight. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Unsupervised neural machine translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018 · 2018
Cited alongside, same era.
The MADAR Arabic dialect corpus and lexicon
Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, and Kemal Oflazer. 2018 · 2018
Cited alongside, same era.
Adapting word embeddings to new languages with morphological and phonological subword representations
Aditi Chaudhary, Chunting Zhou, Lori Levin, Graham Neubig, David R. Mortensen, and Jaime Carbonell. 2018 · 2018
Cited alongside, same era.
Neural machine translation into language varieties
Surafel Melaku Lakew, Aliia Erofeeva, and Marcello Federico. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
CCAligned: A massive collection of cross-lingual web-document pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, and Philipp Koehn. 2020 · 2020
Later among the works it cites.
A multilingual view of unsupervised machine translation
Xavier Garcia, Pierre Foret, Thibault Sellam, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
A probabilistic formulation of unsupervised text style transfer
Junxian He, Xinyi Wang, Graham Neubig, and Taylor Berg-Kirkpatrick. 2020 · 2020
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. 2020 · 2020
Later among the works it cites.
When does unsupervised machine translation work?
Kelly Marchisio, Kevin Duh, and Philipp Koehn. 2020 · 2020
Later among the works it cites.
Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matt Post. 2018 · 2018
Cited alongside, same era.
A unified approach to quantifying algorithmic unfairness: Measuring individual & group unfairness via inequality indices
Till Speicher, Hoda Heidari, Nina Grgic-Hlaca, Krishna P Gummadi, Adish Singla, Adrian Weller, and Muhammad Bilal Zafar. 2018 · 2018
Cited alongside, same era.
Unsupervised text style transfer using language models as discriminators
Zichao Yang, Zhiting Hu, Chris Dyer, Eric P Xing, and Taylor Berg-Kirkpatrick. 2018 · 2018
Cited alongside, same era.
JW300: A wide-coverage parallel corpus for low-resource languages
Željko Agić and Ivan Vulić. 2019 · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample. 2019 · 2019
Cited alongside, same era.
Ethnologue: Languages of the world. 2019
David M Eberhard, Gary F Simons, and Charles D. (eds.) Fennig. 2019 · 2019
Cited alongside, same era.
ParaCrawl: Web-scale parallel corpora for the languages of the EU
Miquel Esplà, Mikel Forcada, Gema Ramírez-Sánchez, and Hieu Hoang. 2019 · 2019
Cited alongside, same era.
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Later among the works it cites.
A monolingual approach to contextualized word embeddings for mid-resource languages
Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2020 · 2020
Later among the works it cites.
Stanza: A Python natural language processing toolkit for many human languages
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020 · 2020
Later among the works it cites.
Making monolingual sentence embeddings multilingual using knowledge distillation
Nils Reimers and Iryna Gurevych. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
BERTscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
How linguistically fair are multilingual pre-trained language models?
Monojit Choudhury and Amit Deshpande. 2021 · 2021
Closest in time.
SD-QA: Spoken Dialectal Question Answering for the Real World
Fahim Faisal, Sharlina Keshava, Md Mahfuz ibn Alam, and Antonios Anastasopoulos. 2021 · 2021
Closest in time.
Harnessing multilinguality in unsupervised machine translation for rare languages
Xavier Garcia, Aditya Siddhant, Orhan Firat, and Ankur Parikh. 2021 · 2021
Closest in time.
An exploration of data augmentation techniques for improving English to Tigrinya translation
Lidia Kidane, Sachin Kumar, and Yulia Tsvetkov. 2021 · 2021
Closest in time.
WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2021 · 2021
Closest in time.