Fetching the paper…
Reading the bibliography…
We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target policy in terms of transition probabilities.
Dynamic Programming
Richard Bellman · 1957
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
On-line q-learning using connectionist systems
Gavin A. Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitry P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Incremental multi-step q-learning
Jing Peng and Ronald J. Williams · 1996
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Richard S. Sutton · 1996
Cited alongside, same era.
Learning to drive a bicycle using reinforcement learning and shaping
Jette Randløv and Preben Alstrøm · 1998
Cited alongside, same era.
Analytical mean squared error curves for temporal difference learning
Satinder Singh and Peter Dayan · 1998
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 1998
Cited alongside, same era.
Bias-variance error bounds for temporal difference updates
Michael J. Kearns and Satinder P. Singh · 2000
Cited alongside, same era.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder Singh · 2000
Cited alongside, same era.
A theoretical and empirical analysis of expected sarsa
Harm van Seijen, Hado van Hasselt, Shimon Whiteson, and Marco Wiering · 2009
Later among the works it cites.
Insights in Reinforcement Learning: formal analysis and empirical evaluation of temporal-difference learning algorithms
Hado Philip van Hasselt · 2011
Later among the works it cites.
A new q (
Richard S. Sutton, Ashique R. Mahmood, Doina Precup, and Hado van Hasselt · 2014
Later among the works it cites.
True online TD(
Harm van Seijen and Richard S. Sutton · 2014
Later among the works it cites.
Generalized emphatic temporal difference learning: Bias-variance analysis
Assaf Hallak, Aviv Tamar, Rémi Munos, and Shie Mannor · 2015
Later among the works it cites.
Off-policy learning based on weighted importance sampling with linear computational complexity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Off-policy temporal-difference learning with function approximation
Doina Precup, Richard S. Sutton, and Sanjoy Dasgupta · 2001
Cited alongside, same era.
Ashique R. Mahmood and Richard S. Sutton · 2015
Later among the works it cites.
Emphatic temporal-difference learning
Ashique R. Mahmood, Huizhen Yu, Martha White, and Richard S. Sutton · 2015
Later among the works it cites.