Understand
Parallel text is required for building high-quality machine translation (MT) systems, as well as for other multilingual NLP applications.
- For many South Asian languages, such data is in short supply.
- In this paper, we described a new publicly available corpus (PMIndia) consisting of parallel sentences which pair 13 major languages of India with English.
- The corpus includes up to 56000 sentences for each language pair.
Built on
WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2019 · 1907
Earlier work this paper cites.
Parallel corpora for medium density languages
D. Varga, L. Németh, P. Halácsy, A. Kornai, V. Trón, and V. Nagy. 2005 · 2005
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014 · 2014
Earlier work this paper cites.
The language demographics of Amazon Mechanical Turk
Ellie Pavlick, Matt Post, Ann Irvine, Dmitry Kachaev, and Chris Callison-Burch. 2014 · 2014
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Similar
Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
Mikel Artetxe and Holger Schwenk. 2018 · 2018
Cited alongside, same era.
The iit bombay english-hindi parallel corpus
Anoop Kunchukuttan, Pratik Mehta, and Pushpak Bhattacharyya. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting bleu scores
Matt Post. 2018 · 2018
Cited alongside, same era.
JW300: A wide-coverage parallel corpus for low-resource languages
Željko Agić and Ivan Vulić. 2019 · 2019
Cited alongside, same era.
Findings of the 2019 conference on machine translation (WMT19)
Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019 · 2019
Cited alongside, same era.
Then
Exploiting multilingualism through multistage fine-tuning for low-resource neural machine translation
Raj Dabre, Atsushi Fujita, and Chenhui Chu. 2019 · 2019
Later among the works it cites.
The FLORES evaluation datasets for low-resource machine translation: Nepali–English and Sinhala–English
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc’Aurelio Ranzato. 2019 · 2019
Later among the works it cites.
Findings of the WMT 2019 shared task on parallel corpus filtering for low-resource conditions
Philipp Koehn, Francisco Guzmán, Vishrav Chaudhary, and Juan Pino. 2019 · 2019
Later among the works it cites.
Revisiting low-resource neural machine translation: A case study
Rico Sennrich and Biao Zhang. 2019 · 2019
Later among the works it cites.
Vecalign: Improved sentence alignment in linear time and space
Brian Thompson and Philipp Koehn. 2019 · 2019
Later among the works it cites.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…