Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
Tesauro, G. et al · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Bertsekas, D. P. and Tsitsiklis, J. N · 1996
Earlier work this paper cites.
A new look at independence
Talagrand, M · 1996
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D · 2000
Earlier work this paper cites.
Gendice: Generalized offline estimation of stationary values
Original
Zhang, R., Dai, B., Li, L., and Schuurmans, D · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M. et al · 2003
Earlier work this paper cites.
Local rademacher complexities
Bartlett, P. L., Bousquet, O., and Mendelson, S · 2005
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Error bounds for approximate value iteration
Munos, R · 2005
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
Batch value-function approximation with only realizability
Original
Xie, T. and Jiang, N · 2008
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Original
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Earlier work this paper cites.