Fetching the paper…
Reading the bibliography…
In this paper, a new offline actor-critic learning algorithm is introduced: Sampled Policy Gradient (SPG).
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Learning from Delayed Rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Adaptive critic designs
Danil V Prokhorov and Donald C Wunsch · 1997
Earlier work this paper cites.
QV ( λ \lambda )-learning: A new on-policy Reinforcement Learning Algorithm
M. Wiering · 2005
Earlier work this paper cites.
Reinforcement Learning in Continuous Action Spaces
H. Van Hasselt and M. A. Wiering · 2007
Cited alongside, same era.
Double Q-learning
Hado V. Hasselt · 2010
Cited alongside, same era.
Connectionist reinforcement learning for intelligent unit micro management in Starcraft
A. Shantia, E. Begue, and M. A. Wiering · 2011
Cited alongside, same era.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2012
Cited alongside, same era.
Playing Atari with Deep Reinforcement Learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Later among the works it cites.
OpenAI baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu · 2017
Later among the works it cites.
Deep Reinforcement Learning that Matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2017
Later among the works it cites.
Playing FPS games with deep reinforcement learning
Guillaume Lample and Devendra Singh Chaplot · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2017
Later among the works it cites.
Opponent modelling in the game of tron using reinforcement learning
Stefan J. L. Knegt, Madalina M. Drugan, and Marco Wiering · 2018
Closest in time.