2020

PMIndia -- A Collection of Parallel Corpora of Languages of India

Haddow, Barry, Kirefu, Faheem

Understand

Parallel text is required for building high-quality machine translation (MT) systems, as well as for other multilingual NLP applications.

  • For many South Asian languages, such data is in short supply.
  • In this paper, we described a new publicly available corpus (PMIndia) consisting of parallel sentences which pair 13 major languages of India with English.
  • The corpus includes up to 56000 sentences for each language pair.

Built on

  • WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia

    Original

    Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2019 · 1907

    Earlier work this paper cites.

  • Parallel corpora for medium density languages

    D. Varga, L. Németh, P. Halácsy, A. Kornai, V. Trón, and V. Nagy. 2005 · 2005

    Earlier work this paper cites.

  • Moses: Open source toolkit for statistical machine translation

    Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007

    Earlier work this paper cites.

  • Findings of the 2014 workshop on statistical machine translation

    Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014 · 2014

    Earlier work this paper cites.

  • The language demographics of Amazon Mechanical Turk

    Ellie Pavlick, Matt Post, Ann Irvine, Dmitry Kachaev, and Chris Callison-Burch. 2014 · 2014

    Earlier work this paper cites.

  • Neural machine translation of rare words with subword units

    Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016

    Earlier work this paper cites.

Similar

  • Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond

    Original

    Mikel Artetxe and Holger Schwenk. 2018 · 2018

    Cited alongside, same era.

  • The iit bombay english-hindi parallel corpus

    Anoop Kunchukuttan, Pratik Mehta, and Pushpak Bhattacharyya. 2018 · 2018

    Cited alongside, same era.

  • A call for clarity in reporting bleu scores

    Matt Post. 2018 · 2018

    Cited alongside, same era.

  • JW300: A wide-coverage parallel corpus for low-resource languages

    Željko Agić and Ivan Vulić. 2019 · 2019

    Cited alongside, same era.

  • Findings of the 2019 conference on machine translation (WMT19)

    Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019 · 2019

    Cited alongside, same era.

Then

  • Exploiting multilingualism through multistage fine-tuning for low-resource neural machine translation

    Raj Dabre, Atsushi Fujita, and Chenhui Chu. 2019 · 2019

    Later among the works it cites.

  • The FLORES evaluation datasets for low-resource machine translation: Nepali–English and Sinhala–English

    Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc’Aurelio Ranzato. 2019 · 2019

    Later among the works it cites.

  • Findings of the WMT 2019 shared task on parallel corpus filtering for low-resource conditions

    Philipp Koehn, Francisco Guzmán, Vishrav Chaudhary, and Juan Pino. 2019 · 2019

    Later among the works it cites.

  • Revisiting low-resource neural machine translation: A case study

    Rico Sennrich and Biao Zhang. 2019 · 2019

    Later among the works it cites.

  • Vecalign: Improved sentence alignment in linear time and space

    Brian Thompson and Philipp Koehn. 2019 · 2019

    Later among the works it cites.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…