2019

Facebook FAIR's WMT19 News Translation Task Submission

Ng, Nathan, Yee, Kyra, Baevski, Alexei et al.

Understand

This paper describes Facebook FAIR's submission to the WMT19 shared news translation task.

  • We participate in two language pairs and four language directions, English <-> German and English <-> Russian.
  • Following our submission from last year, our baseline systems are large BPE-based transformer models trained with the Fairseq sequence modeling toolkit which rely on sampled back-translations.
  • This year we experiment with different bitext data filtering schemes, as well as with adding filtered back-translated data.

Built on

  • Moses: Open source toolkit for statistical machine translation

    Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007

    Earlier work this paper cites.

  • Intelligent selection of language model training data

    Robert Moore and William Lewis. 2010 · 2010

    Earlier work this paper cites.

  • Kenlm: faster and smaller language model queries

    Kenneth Heafield. 2011 · 2011

    Earlier work this paper cites.

  • langid. py: An off-the-shelf language identification tool

    Marco Lui and Timothy Baldwin. 2012 · 2012

    Earlier work this paper cites.

Similar

  • Bag of tricks for efficient text classification

    Original

    Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2016 · 2016

    Cited alongside, same era.

  • Neural machine translation of rare words with subword units

    Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016

    Cited alongside, same era.

  • Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017

    Cited alongside, same era.

  • Understanding back-translation at scale

    Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018 · 2018

    Cited alongside, same era.

Then

  • Scaling neural machine translation

    Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018

    Later among the works it cites.

  • A call for clarity in reporting BLEU scores

    Matt Post. 2018 · 2018

    Later among the works it cites.

  • Findings of the 2019 conference on machine translation (wmt19)

    Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, and Christof Monz. 2019 · 2019

    Closest in time.

  • fairseq: A fast, extensible toolkit for sequence modeling

    Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…