Fetching the paper…
Reading the bibliography…
Most existing deep reinforcement learning (DRL) frameworks consider either discrete action space or continuous action space solely.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Stochastic approximation with two time scales
V. S. Borkar · 1997
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2002
Earlier work this paper cites.
Stochastic Approximation and Recursive Algorithms and Applications
H. Kushner and G. Yin · 2006
Earlier work this paper cites.
Agent2d base code
H. Akiyama · 2010
Earlier work this paper cites.
Neural networks for machine learning-lecture 6a-overview of mini-batch gradient descent, 2012
G. Hinton, N. Srivastava, and K. Swersky · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Continuous deep q q -learning with model-based acceleration
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Guided policy search as approximate mirror descent
W. Montgomery and S. Levine · 2016
Later among the works it cites.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Later among the works it cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas · 2016
Later among the works it cites.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. v. Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Deep reinforcement learning in parameterized action space
M. Hausknecht and P. Stone · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Reinforcement learning with parameterized actions
W. Masson, P. Ranchod, and G. Konidaris · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
O. Anschel, N. Baram, and N. Shimkin · 2017
Later among the works it cites.
Active exploration and parameterized reinforcement learning applied to a simulated human-robot interaction task
M. Khamassi, G. Velentzas, T. Tsitsimis, and C. Tzafestas · 2017
Later among the works it cites.
Pgq: Combining policy gradient and q q -learning
B. O’Donoghue, R. Munos, K. Kavukcuoglu, and V. Mnih · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis · 2017
Later among the works it cites.