Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Original
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign, iwslt 2014
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael J. Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
Recurrent highway networks
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber · 2014
Earlier work this paper cites.
Gated feedback recurrent neural networks
Junyoung Chung, Çaglar Gülçehre, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Hierarchical memory networks
Original
Sarath Chandar, Sungjin Ahn, Hugo Larochelle, Pascal Vincent, Gerald Tesauro, and Yoshua Bengio · 2016
Earlier work this paper cites.
Hierarchical multiscale recurrent neural networks
Original
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2016
Earlier work this paper cites.
Hypernetworks
Original
David Ha, Andrew Dai, and Quoc V Le · 2016
Earlier work this paper cites.
Multiplicative lstm for sequence modelling
Original
Ben Krause, Liang Lu, Iain Murray, and Steve Renals · 2016
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
André F. T. Martins and Ramón Fernández Astudillo · 2016
Earlier work this paper cites.
Scaling memory-augmented neural networks with sparse reads and writes
Jack Rae, Jonathan J Hunt, Ivo Danihelka, Timothy Harley, Andrew W Senior, Gregory Wayne, Alex Graves, and Timothy Lillicrap · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.