Fetching the paper…
Reading the bibliography…
Policy gradient algorithms have found wide adoption in NLP, but have recently become subject to criticism, doubting their suitability for NMT.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams. 1992 · 1992
Earlier work this paper cites.
Deterministic annealing for clustering, compression, classification, regression and related optimization problems
Kenneth Rose. 1998 · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto. 1998 · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Findings of the 2015 workshop on statistical machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Barry Haddow, Matthias Huck, Chris Hokamp, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Carolina Scarton, Lucia Specia, and Marco Turchi. 2015 · 2015
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. 2016 · 2016
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Minimum risk training for neural machine translation
Shiqi Shen, Yong Cheng, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu. 2016 · 2016
Earlier work this paper cites.
Stochastic structured prediction under bandit feedback
Artem Sokolov, Julia Kreutzer, Stefan Riezler, and Christopher Lo. 2016 · 2016
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron C. Courville, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. 2017 · 2017
Cited alongside, same era.
Bandit structured prediction for neural sequence-to-sequence learning
Julia Kreutzer, Artem Sokolov, and Stefan Riezler. 2017 · 2017
Cited alongside, same era.
Reinforcement learning for bandit neural machine translation with simulated human feedback
Khanh Nguyen, Hal Daumé III, and Jordan Boyd-Graber. 2017 · 2017
Cited alongside, same era.
Structured prediction via learning to search under bandit feedback
Amr Sharaf and Hal Daumé III. 2017 · 2017
Cited alongside, same era.
A shared task on bandit learning for machine translation
Artem Sokolov, Julia Kreutzer, Kellen Sunderland, Pavel Danchenko, Witold Szymaniak, Hagen Fürstenau, and Stefan Riezler. 2017 · 2017
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Later among the works it cites.
A study of reinforcement learning for neural machine translation
Lijun Wu, Fei Tian, Tao Qin, Jianhuang Lai, and Tie-Yan Liu. 2018 · 2018
Later among the works it cites.
Breaking the beam search curse: A study of (re-)scoring methods and stopping criteria for neural machine translation
Yilin Yang, Liang Huang, and Mingbo Ma. 2018 · 2018
Later among the works it cites.
Historical text normalization with delayed rewards
Simon Flachs, Marcel Bollmann, and Anders Søgaard. 2019 · 2019
Later among the works it cites.
Joey NMT: A minimalist NMT toolkit for novices
Julia Kreutzer, Jasmijn Bastings, and Stefan Riezler. 2019 · 2019
Later among the works it cites.
Deep reinforcement learning for modeling chit-chat dialog with discrete attributes
Chinnadhurai Sankar and Sujith Ravi. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Adams Wei Yu, Hongrae Lee, and Quoc Le. 2017 · 2017
Cited alongside, same era.
Classical structured prediction losses for sequence to sequence learning
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato. 2018 · 2018
Cited alongside, same era.
Paraphrase generation with deep reinforcement learning
Zichao Li, Xin Jiang, Lifeng Shang, and Hang Li. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Beyond BLEU:training neural machine translation with semantic similarity
John Wieting, Taylor Berg-Kirkpatrick, Kevin Gimpel, and Graham Neubig. 2019 · 2019
Later among the works it cites.
On the weaknesses of reinforcement learning for neural machine translation
Leshem Choshen, Lior Fox, Zohar Aizenbud, and Omri Abend. 2020 · 2020
Later among the works it cites.
On exposure bias, hallucination and domain shift in neural machine translation
Chaojun Wang and Rico Sennrich. 2020 · 2020
Later among the works it cites.