Fetching the paper…
Reading the bibliography…
Previous neural machine translation models used some heuristic search algorithms (e.g., beam search) in order to avoid solving the maximum a posteriori problem over translation sentences at test time.
Statistical theory of extreme values and some practical applications: a series of lectures
Emil Julius Gumbel. 1954 · 1954
Earlier work this paper cites.
Learning representations by back-propagating errors
David Rumelhart, Geoffrey Hinton, and Ronald Williams. 1986 · 1986
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J Williams and David Zipser. 1989 · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
Chin-Yew Lin and Franz Josef Och. 2004 · 2004
Earlier work this paper cites.
Perturb-and-MAP random fields: Using discrete optimization to learn and sample from energy models
George Papandreou and Alan L Yuille. 2011 · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton. 2012 · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler. 2012 · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
A* sampling
Chris J Maddison, Daniel Tarlow, and Tom Minka. 2014 · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Cited alongside, same era.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio. 2016 · 2016
Later among the works it cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2016 · 2016
Later among the works it cites.
GANS for Sequences of Discrete Elements with the Gumbel-softmax Distribution
M. J. Kusner and J. M. Hernández-Lobato. 2016 · 2016
Later among the works it cites.
Professor forcing: A new algorithm for training recurrent networks
Alex M Lamb, Anirudh Goyal ALIAS PARTH GOYAL, Ying Zhang, Saizheng Zhang, Aaron C Courville, and Yoshua Bengio. 2016 · 2016
Later among the works it cites.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh. 2016 · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alec Radford, Luke Metz, and Soumith Chintala. 2015 · 2015
Cited alongside, same era.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2015 · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015 · 2015
Cited alongside, same era.
Minimum risk training for neural machine translation
Shiqi Shen, Yong Cheng, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu. 2015 · 2015
Cited alongside, same era.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Generating images with recurrent adversarial networks
Daniel Jiwoong Im, Chris Dongjoo Kim, Hui Jiang, and Roland Memisevic. 2016a
Cited in the paper.
Generating adversarial parallelization
Daniel Jiwoong Im, He Ma, Chris Dongjoo Kim, and Graham Taylor. 2016b
Cited in the paper.
Later among the works it cites.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M Rush. 2016 · 2016
Later among the works it cites.
Seqgan: sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu. 2016 · 2016
Later among the works it cites.
Energy-based generative adversarial network
Junbo Zhao, Michael Mathieu, and Yann LeCun. 2016 · 2016
Later among the works it cites.
Learning to decode for future success
Jiwei Li, Will Monroe, and Dan Jurafsky. 2017 · 2017
Closest in time.