Fetching the paper…
Reading the bibliography…
We present the first real-world application of methods for improving neural machine translation (NMT) with human reinforcement, based on explicit and implicit user feedback collected on the eBay e-commerce platform.
Measuring nominal scale agreement among many raters
Joseph L Fleiss. 1971 · 1971
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh. 2000 · 2000
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, et al. 2007 · 2007
Earlier work this paper cites.
Exploration scavenging
John Langford, Alexander Strehl, and Jennifer Wortman. 2008 · 2008
Earlier work this paper cites.
News from opus-a collection of multilingual parallel corpora with tools and interfaces
Jörg Tiedemann. 2009 · 2009
Earlier work this paper cites.
Learning from logged implicit exploration data
Alexander L. Strehl, John Langford, Lihong Li, and Sham M. Kakade. 2010 · 2010
Earlier work this paper cites.
Better hypothesis testing for statistical machine translation: Controlling for optimizer instability
Jonathan H. Clark, Chris Dyer, Alon Lavie, and Noah A. Smith. 2011 · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li. 2011 · 2011
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X. Charles, D. Max Chickering, Elon Portugaly, Dipanakar Ray, Patrice Simard, and Ed Snelson. 2013 · 2013
Cited alongside, same era.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2015 · 2015
Cited alongside, same era.
Minimum risk training for neural machine translation
Shiqi Shen, Yong Cheng, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu. 2016 · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Later among the works it cites.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio. 2017 · 2017
Later among the works it cites.
Using images to improve machine-translating e-commerce product listings
Iacer Calixto, Daniel Stein, Evgeny Matusov, Pintu Lohar, Sheila Castilho, and Andy Way. 2017 · 2017
Later among the works it cites.
Bandit structured prediction for neural sequence-to-sequence learning
Julia Kreutzer, Artem Sokolov, and Stefan Riezler. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Stanford neural machine translation systems for spoken language domains
Minh-Thang Luong and Christopher D Manning. 2015 · 2015
Cited alongside, same era.
The self-normalized estimator for counterfactual learning
Adith Swaminathan and Thorsten Joachims. 2015 · 2015
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Y Gal and Z Ghahramani. 2016 · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li. 2016 · 2016
Cited alongside, same era.
Counterfactual learning for machine translation: Degeneracies and solutions
Carolin Lawrence, Pratik Gajane, and Stefan Riezler. 2017a
Cited in the paper.
Edinburgh neural machine translation systems for wmt 16
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016a
Cited in the paper.
Counterfactual learning from bandit feedback under deterministic logging : A case study in statistical machine translation
Carolin Lawrence, Artem Sokolov, and Stefan Riezler. 2017b · 2017
Later among the works it cites.
Do convolutional networks need to be deep for text classification ?
Hoa T. Le, Christophe Cerisara, and Alexandre Denis. 2017 · 2017
Later among the works it cites.
Reinforcement learning for bandit neural machine translation with simulated human feedback
Khanh Nguyen, Hal Daumé III, and Jordan Boyd-Graber. 2017 · 2017
Later among the works it cites.
A shared task on bandit learning for machine translation
Artem Sokolov, Julia Kreutzer, Kellen Sunderland, Pavel Danchenko, Witold Szymaniak, Hagen Fürstenau, and Stefan Riezler. 2017 · 2017
Later among the works it cites.