Fetching the paper…
Reading the bibliography…
We propose a novel decoding approach for neural machine translation (NMT) based on continuous optimisation.
A maximum entropy approach to natural language processing
Adam L. Berger, Vincent J. Della Pietra, and Stephen A. Della Pietra. 1996 · 1996
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jurgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Exponentiated Gradient Versus Gradient Descent for Linear Predictors
Jyrki Kivinen and Manfred K. Warmuth. 1997 · 1997
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian. 1999 · 1999
Earlier work this paper cites.
Discriminative training and maximum entropy models for statistical machine translation
Franz Josef Och and Hermann Ney. 2002 · 2002
Earlier work this paper cites.
BLEU: A Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Word-based alignment, phrase-based translation: What’s the link
Adam Lopez and Philip Resnik. 2006 · 2006
Earlier work this paper cites.
Exponentiated Gradient Algorithms for Log-linear Structured Prediction
Amir Globerson, Terry Y. Koo, Xavier Carreras, and Michael Collins. 2007 · 2007
Earlier work this paper cites.
Approximate Inference in Graphical Models using LP Relaxations
David Sontag. 2010 · 2010
Earlier work this paper cites.
Generating Sequences With Recurrent Neural Networks
A. Graves. 2013 · 2013
Earlier work this paper cites.
Report on the 11th IWSLT Evaluation Campaign
M. Cettolo, J. Niehues, S. St¨uker, L. Bentivogli, and M. Federico. 2014 · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Cited alongside, same era.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
A Critical Review of Recurrent Neural Networks for Sequence Learning
Z. C. Lipton, J. Berkowitz, and C. Elkan. 2015 · 2015
Cited alongside, same era.
Addressing the Rare Word Problem in Neural Machine Translation
Thang Luong, Ilya Sutskever, Quoc Le, Oriol Vinyals, and Wojciech Zaremba. 2015 · 2015
Cited alongside, same era.
Structured prediction energy networks
David Belanger and Andrew McCallum. 2016 · 2016
Cited alongside, same era.
Mutual Information and Diverse Decoding Improve Neural Machine Translation
J. Li and D. Jurafsky. 2016 · 2016
Later among the works it cites.
A Simple, Fast Diverse Decoding Algorithm for Neural Generation
J. Li, W. Monroe, and D. Jurafsky. 2016 · 2016
Later among the works it cites.
Coverage Embedding Models for Neural Machine Translation
Haitao Mi, Baskaran Sankaran, Zhiguo Wang, and Abe Ittycheriah. 2016 · 2016
Later among the works it cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Later among the works it cites.
Decoding neural machine translation using gradient descent
Emanuel Snelleman. 2016 · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Edinburgh Neural Machine Translation Systems for WMT 16
Rico Sennrich; Barry Haddow; Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Noisy Parallel Approximate Decoding for Conditional Recurrent Language Model
K. Cho. 2016 · 2016
Cited alongside, same era.
Incorporating Structural Alignment Biases into an Attentional Neural Translation Model
Trevor Cohn, Cong Duy Vu Hoang, Ekaterina Vymolova, Kaisheng Yao, Chris Dyer, and Gholamreza Haffari. 2016 · 2016
Cited alongside, same era.
Bidirectional Decoding for Statistical Machine Translation
Taro Watanabe and Eiichiro Sumita. 2002a
Cited in the paper.
Bidirectional decoding for statistical machine translation
Taro Watanabe and Eiichiro Sumita. 2002b
Cited in the paper.
Sam Wiseman and Alexander M. Rush. 2016 · 2016
Later among the works it cites.
End-to-End Learning for Structured Prediction Energy Networks
D. Belanger, B. Yang, and A. McCallum. 2017 · 2017
Closest in time.
Differentiable scheduled sampling for credit assignment
Kartik Goyal, Chris Dyer, and Taylor Berg-Kirkpatrick. 2017 · 2017
Closest in time.
DyNet: The Dynamic Neural Network Toolkit
G. Neubig, C. Dyer, Y. Goldberg, A. Matthews, W. Ammar, A. Anastasopoulos, M. Ballesteros, D. Chiang, D. Clothiaux, T. Cohn, K. Duh, M. Faruqui, C. Gan, D. Garrette, Y. Ji, L. Kong, A. Kuncoro, G. Kumar, C. Malaviya, P. Michel, Y. Oda, M. Richardson, N. Saphra, S. Swayamdipta, and P. Yin. 2017 · 2017
Closest in time.