Fetching the paper…
Reading the bibliography…
In this work, we take a fresh look at some old and new algorithms for off-policy, return-based reinforcement learning.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
Watkins, C. J. C. H. (1989) · 1989
Earlier work this paper cites.
Scaling up reinforcement learning for robot control
Lin, L. (1993) · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Sutton, R. S. (1996) · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. and Barto, A. (1998) · 1998
Earlier work this paper cites.
Bias-variance error bounds for temporal difference updates
Kearns, M. J. and Singh, S. P. (2000) · 2000
Cited alongside, same era.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R. S., and Singh, S. (2000) · 2000
Cited alongside, same era.
Convergence results for single-step on-policy reinforcement-learning algorithms
Singh, S., Jaakkola, T., Littman, M. L., and Szepesvári, C. (2000) · 2000
Cited alongside, same era.
Off-policy temporal-difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S. (2001) · 2001
Cited alongside, same era.
On the convergence of optimistic policy iteration
Tsitsiklis, J. N. (2003) · 2003
Cited alongside, same era.
The Arcade Learning Environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Generalized emphatic temporal difference learning: Bias-variance analysis
Hallak, A., Tamar, A., Munos, R., and Mannor, S. (2015) · 2015
Later among the works it cites.
Off-policy learning based on weighted importance sampling with linear computational complexity
Mahmood, A. R. and Sutton, R. S. (2015) · 2015
Later among the works it cites.
Emphatic temporal-difference learning
Mahmood, A. R., Yu, H., White, M., and Sutton, R. S. (2015) · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Later among the works it cites.
Q( λ \lambda ) with off-policy corrections
Harutyunyan, A., Bellemare, M. G., Stepleton, T., and Munos, R. (2016) · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Off-policy learning with eligibility traces: A survey
Geist, M. and Scherrer, B. (2014) · 2014
Cited alongside, same era.
Closest in time.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 2016
Closest in time.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2016) · 2016
Closest in time.