Fetching the paper…
Reading the bibliography…
Off-policy learning refers to the problem of learning the value function of a way of behaving, or policy, while following a different policy.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Q-learning
ChristopherJ.C.H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Chattering in sarsa(lambda) - a cmu learning lab internal report
Geoffrey J. Gordon · 1996
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Cited alongside, same era.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 2000
Cited alongside, same era.
Reinforcement learning with function approximation converges to a region
Geoffrey J. Gordon · 2001
Cited alongside, same era.
A convergent form of approximate policy iteration
Theodore J. Perkins and Doina Precup · 2003
Cited alongside, same era.
Gq(lambda): A general gradient algorithm for temporal-difference prediction learning with eligibility traces
Sutton R. S. Maei, H. R
Cited in the paper.
Toward off-policy learning control with function approximation
Szepesvari Cs. Bhatnagar S. Sutton R. S. Maei, H. R
Cited in the paper.
A convergent o(n) algorithm for off-policy temporal-difference learning with linear function approximation
Richard S. Sutton, Csaba Szepesvári, and Hamid Reza Maei
Cited in the paper.
Convergent temporal-difference learning with arbitrary smooth function approximation
Szepesvari Cs. Bhatnagar S. Precup D. Silver D. Sutton R. S. Maei, H. R · 2009
Later among the works it cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S. Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Later among the works it cites.
Policy evaluation with temporal differences: A survey and comparison
Christoph Dann, Gerhard Neumann, and Jan Peters · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…