Fetching the paper…
Reading the bibliography…
A deep learning approach to reinforcement learning led to a general learner able to train on visual input to play a variety of arcade games at the human and superhuman levels.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jurgen Schmidhuber · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David Mcallester, Satinder Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Learning where to attend with deep architectures for image tracking
Misha Denil, Loris Bazzani, Hugo Larochelle, and Nando de Freitas · 2012
Earlier work this paper cites.
Recurrent models of visual attention
Volodymyr Mnih, Nicolas Heess, Alex Graves, and koray kavukcuoglu · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Deep recurrent q-learning for partially observable mdps
Matthew J. Hausknecht and Peter Stone · 2015
Cited alongside, same era.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, et al · 2015
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio · 2015
Closest in time.
Multiple object recognition with visual attention
Jimmy Ba, Volodymyr Mnih, and Koray Kavukcuoglu · 2015
Closest in time.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Closest in time.
Learning wake-sleep recurrent attention models
Jimmy Ba, Roger Grosse, Ruslan Salakhutdinov, and Brendan Frey · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Nicolas Heess, Theophane Weber, and Pieter Abbeel · 2015
Closest in time.