Fetching the paper…
Reading the bibliography…
Sequence-to-sequence models with soft attention had significant success in machine translation, speech recognition, and question answering.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
The hitch hiker’s guide to the galaxy: a trilogy in five parts
Douglas Adams · 1995
Earlier work this paper cites.
Acoustic modeling using deep belief networks
Abdel-rahman Mohamed, George E. Dahl, and Geoffrey Hinton · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alan Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwen, and Yoshua Bengio · 2014
Earlier work this paper cites.
End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
Jan Chorowski, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Towards End-to-End Speech Recognition with Recurrent Neural Networks
Alex Graves and Navdeep Jaitly · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Motor Skill Learning with Local Trajectory Methods
Sergey Levine · 2014
Cited alongside, same era.
Neural variational inference and learning in belief networks
Andriy Mnih and Karol Gregor · 2014
Cited alongside, same era.
Recurrent models of visual attention
Volodymyr Mnih, Nicolas Heess, Alex Graves, et al · 2014
Cited alongside, same era.
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Cited alongside, same era.
Jason Weston, Sumit Chopra, and Antoine Bordes · 2014
Cited alongside, same era.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
An online sequence-to-sequence model using partial conditioning
Navdeep Jaitly, Quoc V Le, Oriol Vinyals, Ilya Sutskeyver, and Samy Bengio · 2015
Later among the works it cites.
Nal Kalchbrenner, Ivo Danihelka, and Alex Graves · 2015
Later among the works it cites.
Gradient estimation using stochastic computation graphs
John Schulman, Nicolas Heess, Theophane Weber, and Pieter Abbeel · 2015
Later among the works it cites.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al · 2015
Later among the works it cites.
Show and Tell: A Neural Image Caption Generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end attention-based large vocabulary speech recognition
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio · 2015
Cited alongside, same era.
William Chan, Navdeep Jaitly, Quoc V Le, and Oriol Vinyals · 2015
Cited alongside, same era.
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio · 2015
Later among the works it cites.
Reinforcement learning neural turing machines
Wojciech Zaremba and Ilya Sutskever · 2015
Later among the works it cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2016
Closest in time.