Fetching the paper…
Reading the bibliography…
Using deep neural nets as function approximator for reinforcement learning tasks have recently been shown to be very powerful for solving problems approaching real-world complexity.
Applied dynamic programming
Richard Ernest Bellman and Stuart E Dreyfus · 1962
Earlier work this paper cites.
Cognitive and attentional mechanisms in delay of gratification
Walter Mischel, Ebbe B Ebbesen, and Antonette Raskoff Zeiss · 1972
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Approximate solutions to markov decision processes
Geoffrey J Gordon · 1999
Cited alongside, same era.
A sparse sampling algorithm for near-optimal planning in large markov decision processes
Michael Kearns, Yishay Mansour, and Andrew Y Ng · 2002
Cited alongside, same era.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Cited alongside, same era.
Reinforcement learning and dynamic programming using function approximators
Lucian Busoniu, Robert Babuska, Bart De Schutter, and Damien Ernst · 2010
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2012
Cited alongside, same era.
The dependence of effective planning horizon on model accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis · 2015
Closest in time.
Deep reinforcement learning with double Q-learning
Hado Van Hasselt, Arthur Guez, and Silver David · 2015
Closest in time.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, et al · 2015
Closest in time.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Closest in time.
Dueling network architectures for deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Ziyu Wang, Nando de Freitas, and Marc Lanctot · 2015
Closest in time.
Deep recurrent Q-learning for partially observable MDPs
Matthew Hausknecht and Peter Stone · 2015
Closest in time.