Fetching the paper…
Reading the bibliography…
Multi-step temporal-difference (TD) learning, where the update targets contain information from multiple time steps ahead, is one of the most popular forms of TD learning for linear function approximation.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W · 1983
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G · 1994
Earlier work this paper cites.
Nonlinear Programming
Bertsekas, D. P · 1995
Earlier work this paper cites.
Truncating temporal differences: On the efficient implementation of TD( λ \lambda ) for reinforcement learning
Cichosz, P · 1995
Cited alongside, same era.
Neuro-Dynamic Programming
Bertsekas, D. and Tsitsiklis, J · 1996
Cited alongside, same era.
Reinforcement learning with replacing eligibility traces
Singh, S. P. and Sutton, R. S · 1996
Cited alongside, same era.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 1998
Cited alongside, same era.
Algorithms for Reinforcement Learning
Szepesvári, C · 2009
Later among the works it cites.
True online TD( λ \lambda )
van Seijen, H. and Sutton, R. S · 2014
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., Kumaran, H. King D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Later among the works it cites.
True online temporal-difference learning, 2015
van Seijen, H., Mahmood, A. R., Pilarski, P. M., Machado, M. C., and Sutton, R. S · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…