Fetching the paper…
Reading the bibliography…
We present the first deep learning model to successfully learn control policies directly from high-dimensional sensory input using reinforcement learning.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
Long-Ji Lin · 1993
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less real time
Andrew Moore and Chris Atkeson · 1993
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Why did td-gammon work
Jordan B. Pollack and Alan D. Blair · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto · 1998
Earlier work this paper cites.
Reinforcement learning with factored states and actions
Brian Sallans and Geoffrey E. Hinton · 2004
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Cited alongside, same era.
What is the best multi-stage architecture for object recognition?
Kevin Jarrett, Koray Kavukcuoglu, Marc’Aurelio Ranzato, and Yann LeCun · 2009
Cited alongside, same era.
Convergent Temporal-Difference Learning with Arbitrary Smooth Function Approximation
Hamid Maei, Csaba Szepesvari, Shalabh Bhatnagar, Doina Precup, David Silver, and Rich Sutton · 2009
Cited alongside, same era.
Deep auto-encoder neural networks in reinforcement learning
Sascha Lange and Martin Riedmiller · 2010
Cited alongside, same era.
Toward off-policy learning control with function approximation
Hamid Maei, Csaba Szepesvári, Shalabh Bhatnagar, and Richard S. Sutton · 2010
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
Actor-critic reinforcement learning with energy-based policies
Nicolas Heess, David Silver, and Yee Whye Teh · 2012
Later among the works it cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoff Hinton · 2012
Later among the works it cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Closest in time.
Bayesian learning of recursively factored environments
Marc G. Bellemare, Joel Veness, and Michael Bowling · 2013
Closest in time.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey E. Hinton · 2013
Closest in time.
A neuro-evolution approach to general atari game playing
Matthew Hausknecht, Risto Miikkulainen, and Peter Stone · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vinod Nair and Geoffrey E Hinton · 2010
Cited alongside, same era.
Sketch-based linear value function approximation
Marc Bellemare, Joel Veness, and Michael Bowling · 2012
Cited alongside, same era.
Investigating contingency awareness using atari 2600 games
Marc G Bellemare, Joel Veness, and Michael Bowling · 2012
Cited alongside, same era.
Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition
George E. Dahl, Dong Yu, Li Deng, and Alex Acero · 2012
Cited alongside, same era.
Closest in time.
Machine Learning for Aerial Image Labeling
Volodymyr Mnih · 2013
Closest in time.
Pedestrian detection with unsupervised multi-stage feature learning
Pierre Sermanet, Koray Kavukcuoglu, Soumith Chintala, and Yann LeCun · 2013
Closest in time.