Fetching the paper…
Reading the bibliography…
We describe a mechanism for subsampling sequences and show how to compute its expected output so that it can be trained with standard backpropagation.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Wojciech Zaremba and Ilya Sutskever · 2014
Cited alongside, same era.
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2016
Later among the works it cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Later among the works it cites.
Learning online alignments with continuous rewards policy gradient
Yuping Luo, Chung-Cheng Chiu, Navdeep Jaitly, and Ilya Sutskever · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…