Fetching the paper…
Reading the bibliography…
The temporal-difference methods TD($\lambda$) and Sarsa($\lambda$) form a core part of modern reinforcement learning.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
Watkins, C. J. C. H. (1989) · 1989
Earlier work this paper cites.
The convergence of TD( λ \lambda ) for general λ \lambda
Dayan, P. (1992) · 1992
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, L. P., Littman, M. L., and Moore, A. P. (1996) · 1996
Earlier work this paper cites.
On the worst-case analysis of temporal-difference learning algorithms
Schapire, R. E. and Warmuth, M. K. (1996) · 1996
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Myoelectric signal processing for control of powered limb prostheses
Parker, P., Englehart, K. B., , and Hudgins, B. (2006) · 2006
Earlier work this paper cites.
Algorithms for Reinforcement Learning
Szepesvári, C. (2010) · 2010
Cited alongside, same era.
Gradient Temporal-Difference Learning Algorithms
Maei, H. R. (2011) · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D. (2011) · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
Adaptive artificial limbs: A real-time approach to prediction and anticipation
Pilarski, P. M., Dawson, M. R., Degris, T., Carey, J. P., Chan, K. M., Hebert, J. S., and Sutton, R. S. (2013) · 2013
Cited alongside, same era.
A comparison of learning algorithms on the arcade learning environment
A new Q( λ \lambda ) with interim forward view and Monte Carlo equivalence
Sutton, R. S., Mahmood, A. R., Precup, D., and van Hasselt, H. (2014) · 2014
Later among the works it cites.
Off-policy TD( λ \lambda ) with a true online equivalence
van Hasselt, H., Mahmood, A. R. and Sutton, R. S. (2014) · 2014
Later among the works it cites.
True online TD( λ \lambda )
van Seijen, H. H. and Sutton, R. S. (2014) · 2014
Later among the works it cites.
Off-policy learning based on weighted importance sampling with linear computational complexity
Mahmood, A. R. and Sutton, R. S. (2015) · 2015
Closest in time.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., Kumaran, H. K. D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Closest in time.
Policy evaluation using the Ω \Omega -return
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Defazio, A. and Graepel, T. (2014) · 2014
Cited alongside, same era.
Novel targeted sensory reinnervation technique to restore functional hand sensation after transhumeral amputation
Hebert, J. S., Olson, J. L., Morhart, M. J., Dawson, M. R., Marasco, P. D., Kuiken, T. A., and Chan, K. M. (2014) · 2014
Cited alongside, same era.
Multi-timescale nexting in a reinforcement learning robot
Modayil, J., White, A., and Sutton, R. S. (2014) · 2014
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E. (2009a)
Cited in the paper.
A convergent 𝒪 ( n ) \mathcal{O}(n) algorithm for off-policy temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., and Szepesvári, C. (2009b)
Cited in the paper.
Thomas, P. S., Niekum, S., Theocharous, G., and Konidaris, G. (2015) · 2015
Closest in time.
Learning to predict independent of span
van Hasselt, H. and Sutton, R. S. (2015) · 2015
Closest in time.
Effective multi-step temporal-difference learning for non-linear function approximation
van Seijen, H. H. (2016) · 2016
Closest in time.