Fetching the paper…
Reading the bibliography…
Reinforcement Learning algorithms can learn complex behavioral patterns for sequential decision making tasks wherein an agent interacts with an environment and acquires feedback in the form of rewards sampled from it.
Transition point dynamic programming
Kenneth M Buckland and Peter D Lawrence · 1994
Earlier work this paper cites.
Reinforcement learning methods for continuous-time markov decision problems
Steven J Duff · 1995
Earlier work this paper cites.
Self-improving factory simulation using continuous-time average-reward reinforcement learning
Sridhar Mahadevan, Nicholas Marchalleck, Tapas K Das, and Abhijit Gosavi · 1997
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
Guy Lever · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Cited alongside, same era.
Deep reinforcement learning in parametrized action space
Matthew Hausknecht and Peter Stone · 2016
Later among the works it cites.
Half field offense: An environment for multiagent learning and ad hoc teamwork
Matthew Hausknecht, Prannoy Mupparaju, Sandeep Subramanian, Shivaram Kalyanakrishnan, and Peter Stone · 2016
Later among the works it cites.
Dynamic frame skip deep q network
Aravind S Lakshminarayanan, Sahil Sharma, and Balaraman Ravindran · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Simultaneous machine translation using deep reinforcement learning
Harsh Satija and Joelle Pineau · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Sergey Levine, Philipp Moritz, Michael I Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Deep reinforcement learning with macro-actions
Ishan P Durugkar, Clemens Rosenbaum, Stefan Dernbach, and Sridhar Mahadevan · 2016
Cited alongside, same era.
Learning to translate in real-time with neural machine translation
Jiatao Gu, Graham Neubig, Kyunghyun Cho, and Victor OK Li · 2016
Cited alongside, same era.
Alexander Vezhnevets, Volodymyr Mnih, Simon Osindero, Alex Graves, Oriol Vinyals, John Agapiou, et al · 2016
Later among the works it cites.
Torcs, the open racing car simulator
Bernhard Wymann, E Espié, C Guionneau, C Dimitrakakis, R Coulom, and A Sumner · 2016
Later among the works it cites.