Understand
While neural machine translation (NMT) is making good progress in the past two years, tens of millions of bilingual sentence pairs are needed for its training.
- However, human labeling is very costly.
- To tackle this training data bottleneck, we develop a dual-learning mechanism, which can enable an NMT system to automatically learn from unlabeled data through a dual-learning game.
- This mechanism is inspired by the following observation: any machine translation task has a dual task, e.g., English-to-French translation (primal) versus French-to-English translation (dual); the primal and dual tasks can form a closed loop, and generate informative feedback signals to train the translation models, even if without the involvement of a human labeler.
Built on
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Statistical phrase-based translation
P. Koehn, F. J. Och, and D. Marcu · 2003
Earlier work this paper cites.
Large language models in machine translation
T. Brants, A. C. Popat, P. Xu, F. J. Och, and J. Dean · 2007
Earlier work this paper cites.
Semi-supervised model adaptation for statistical machine translation
N. Ueffing, G. Haffari, and A. Sarkar · 2008
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Similar
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Cited alongside, same era.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
On using monolingual corpora in neural machine translation
C. Gulcehre, O. Firat, K. Xu, K. Cho, L. Barrault, H.-C. Lin, F. Bougares, H. Schwenk, and Y. Bengio · 2015
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
S. Jean, K. Cho, R. Memisevic, and Y. Bengio · 2015
Cited alongside, same era.
Then
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2015
Later among the works it cites.
A neural attention model for abstractive sentence summarization
A. M. Rush, S. Chopra, and J. Weston · 2015
Later among the works it cites.
Agreement-based joint training for bidirectional attention-based neural machine translation
Y. Cheng, S. Shen, Z. He, W. He, H. Wu, M. Sun, and Y. Liu · 2016
Closest in time.
Improving neural machine translation models with monolingual data
R. Sennrich, B. Haddow, and A. Birch · 2016
Closest in time.
Minimum risk training for neural machine translation
S. Shen, Y. Cheng, Z. He, W. He, H. Wu, M. Sun, and Y. Liu · 2016
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…