Fetching the paper…
Reading the bibliography…
Recent advances in conditional recurrent language modelling have mainly focused on network architectures (e.g., attention mechanism), learning algorithms (e.g., scheduled sampling and sequence-level training) and novel applications (e.g., image/video description generation, speech recognition, etc.) On the other hand, we notice that decoding algorithms/strategies have not been investigated as much, and it has become standard to use greedy or beam search.
Learning representations by back-propagating errors
D. Rumelhart, G. Hinton, and R. Williams · 1986
Earlier work this paper cites.
Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition
J. S. Bridle · 1990
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Foundations of statistical natural language processing
C. D. Manning and H. Schütze · 1999
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, and P. Vincent · 2003
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Perturb-and-MAP random fields: Using discrete optimization to learn and sample from energy models
G. Papandreou and A. L. Yuille · 2011
Earlier work this paper cites.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. E. Hinton · 2011
Earlier work this paper cites.
Diverse m-best solutions in Markov random fields
D. Batra, P. Yadollahpour, A. Guzman-Rivera, and G. Shakhnarovich · 2012
Earlier work this paper cites.
Statistical Language Models based on Neural Networks
T. Mikolov · 2012
Cited alongside, same era.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Cited alongside, same era.
Better mixing via deep representations
Y. Bengio, G. Mesnil, Y. Dauphin, and S. Rifai · 2013
Cited alongside, same era.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
Perturb-and-MAP random fields: Reducing random sampling to optimization, with applications in computer vision
Describing multimedia content using attention-based encoder-decoder networks
K. Cho, A. Courville, and Y. Bengio · 2015
Later among the works it cites.
A recurrent latent variable model for sequential data
J. Chung, K. Kastner, L. Dinh, K. Goel, A. Courville, and Y. Bengio · 2015
Later among the works it cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Later among the works it cites.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2015
Later among the works it cites.
Neural machine translation of rare words with subword units
R. Sennrich, B. Haddow, and A. Birch · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Papandreou and A. Yuille · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
Task loss estimation for sequence prediction
D. Bahdanau, D. Serdyuk, P. Brakel, N. R. Ke, J. Chorowski, A. Courville, and Y. Bengio · 2015
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Cited alongside, same era.
M. Sundermeyer, H. Ney, and R. Schlüter · 2015
Later among the works it cites.
Multi-way, multilingual neural machine translation with a shared attention mechanism
O. Firat, K. Cho, and Y. Bengio · 2016
Closest in time.
Mutual information and diverse decoding improve neural machine translation
J. Li and D. Jurafsky · 2016
Closest in time.
Simple, fast noise-contrastive estimation for large RNN vocabularies
B. Zoph, A. Vaswani, J. May, and K. Knight · 2016
Closest in time.