Fetching the paper…
Reading the bibliography…
We know from reinforcement learning theory that temporal difference learning can fail in certain cases.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Dynamic programming and Markov processes
R. A. Howard · 1960
Earlier work this paper cites.
An adaptive optimal controller for discrete-time Markov environments
I. H. Witten · 1977
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
A. G. Barto, R. S. Sutton, and C. W. Anderson · 1983
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Learning from delayed rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L. Lin · 1992
Earlier work this paper cites.
Q-learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
L. Baird · 1995
Earlier work this paper cites.
On the virtues of linear learning and trajectory distributions
R. S. Sutton · 1995
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Introduction to reinforcement learning
R. S. Sutton and A. G. Barto · 1998
Cited alongside, same era.
Off-policy temporal-difference learning with function approximation
D. Precup and R. S. Sutton · 2001
Cited alongside, same era.
Natural actor-critic
J. Peters and S. Schaal · 2008
Cited alongside, same era.
A convergent O(n) algorithm for off-policy temporal-difference learning with linear function approximation
R. S. Sutton, C. Szepesvári, and H. R. Maei · 2008
Cited alongside, same era.
Natural actor-critic algorithms
S. Bhatnagar, R. S. Sutton, M. Ghavamzadeh, and M. Lee · 2009
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
H. R. Maei, C. Szepesvári, S. Bhatnagar, D. Precup, D. Silver, and R. Sutton · 2009
Cited alongside, same era.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Later among the works it cites.
A method for stochastic optimization
D. P. Kingma and J. B. Adam · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Later among the works it cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Double Q-learning
H. van Hasselt · 2010
Cited alongside, same era.
Gradient temporal-difference learning algorithms
H. R. Maei · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Cited alongside, same era.
Off-policy actor-critic
T. Degris, M. White, and R. S. Sutton · 2012
Cited alongside, same era.
Reinforcement learning in continuous state and action spaces
H. van Hasselt · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
R. S. Sutton, A. R. Mahmood, and M. White · 2016
Later among the works it cites.
Deep reinforcement learning with Double Q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, N. de Freitas, T. Schaul, M. Hessel, H. van Hasselt, and M. Lanctot · 2016
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Closest in time.
Distributed prioritized experience replay
D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. van Hasselt, and D. Silver · 2018
Closest in time.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Closest in time.