2016

Dual Learning for Machine Translation

Xia, Yingce, He, Di, Qin, Tao et al.

Understand

While neural machine translation (NMT) is making good progress in the past two years, tens of millions of bilingual sentence pairs are needed for its training.

  • However, human labeling is very costly.
  • To tackle this training data bottleneck, we develop a dual-learning mechanism, which can enable an NMT system to automatically learn from unlabeled data through a dual-learning game.
  • This mechanism is inspired by the following observation: any machine translation task has a dual task, e.g., English-to-French translation (primal) versus French-to-English translation (dual); the primal and dual tasks can form a closed loop, and generate informative feedback signals to train the translation models, even if without the involvement of a human labeler.

Built on

  • Policy gradient methods for reinforcement learning with function approximation

    R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al · 1999

    Earlier work this paper cites.

  • Bleu: a method for automatic evaluation of machine translation

    K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002

    Earlier work this paper cites.

  • Statistical phrase-based translation

    P. Koehn, F. J. Och, and D. Marcu · 2003

    Earlier work this paper cites.

  • Large language models in machine translation

    T. Brants, A. C. Popat, P. Xu, F. J. Och, and J. Dean · 2007

    Earlier work this paper cites.

  • Semi-supervised model adaptation for statistical machine translation

    N. Ueffing, G. Haffari, and A. Sarkar · 2008

    Earlier work this paper cites.

  • Recurrent neural network based language model

    T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010

    Earlier work this paper cites.

Similar

  • Adadelta: an adaptive learning rate method

    Original

    M. D. Zeiler · 2012

    Cited alongside, same era.

  • Learning phrase representations using rnn encoder–decoder for statistical machine translation

    K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014

    Cited alongside, same era.

  • Sequence to sequence learning with neural networks

    I. Sutskever, O. Vinyals, and Q. V. Le · 2014

    Cited alongside, same era.

  • Neural machine translation by jointly learning to align and translate

    D. Bahdanau, K. Cho, and Y. Bengio · 2015

    Cited alongside, same era.

  • On using monolingual corpora in neural machine translation

    Original

    C. Gulcehre, O. Firat, K. Xu, K. Cho, L. Barrault, H.-C. Lin, F. Bougares, H. Schwenk, and Y. Bengio · 2015

    Cited alongside, same era.

  • On using very large target vocabulary for neural machine translation

    S. Jean, K. Cho, R. Memisevic, and Y. Bengio · 2015

    Cited alongside, same era.

Then

  • Sequence level training with recurrent neural networks

    Original

    M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2015

    Later among the works it cites.

  • A neural attention model for abstractive sentence summarization

    A. M. Rush, S. Chopra, and J. Weston · 2015

    Later among the works it cites.

  • Agreement-based joint training for bidirectional attention-based neural machine translation

    Y. Cheng, S. Shen, Z. He, W. He, H. Wu, M. Sun, and Y. Liu · 2016

    Closest in time.

  • Improving neural machine translation models with monolingual data

    R. Sennrich, B. Haddow, and A. Birch · 2016

    Closest in time.

  • Minimum risk training for neural machine translation

    S. Shen, Y. Cheng, Z. He, W. He, H. Wu, M. Sun, and Y. Liu · 2016

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…