Fetching the paper…
Reading the bibliography…
Temporal Difference learning or TD($\lambda$) is a fundamental algorithm in the field of reinforcement learning.
Adjustment of an inverse matrix corresponding to a change in one element of a given matrix
Sherman, J. and Morrison, W. J. (1949) · 1949
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. and Tsitsiklis, J. (1996) · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G. (1996) · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Technical update: Least-squares temporal difference learning
Boyan, J. A. (2002) · 2002
Earlier work this paper cites.
Efficient reinforcement learning using recursive least-squares methods
Xu, X., He, H.-g., and Hu, D. (2002) · 2002
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R. (2003) · 2003
Cited alongside, same era.
Covariate shift adaptation by importance weighted cross validation
Sugiyama, M., Krauledat, M., and MÞller, K.-R. (2007) · 2007
Cited alongside, same era.
Temporal difference bayesian model averaging: A bayesian perspective on adapting lambda
Downey, C. and Sanner, S. (2010) · 2010
Cited alongside, same era.
TD γ \textrm{TD}_{\gamma} : Re-evaluating complex backups in temporal difference learning
Konidaris, G., Niekum, S., and Thomas, P. (2011) · 2011
Cited alongside, same era.
Evaluating recommendation systems
Shani, G. and Gunawardana, A. (2011) · 2011
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Later among the works it cites.
On the Rate of Convergence and Error Bounds for LSTD( λ \lambda )
Tagorti, M. and Scherrer, B. (2015) · 2015
Later among the works it cites.
Policy Evaluation using the Ω \Omega -Return
Thomas, P., Niekum, S., Theocharous, G., and Konidaris, G. (2015) · 2015
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E. (2016) · 2016
Closest in time.
A greedy approach to adapting the trace parameter for temporal difference learning
White, M. and White, A. (2016) · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…