Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Learning to act using real-time dynamic programming
Barto, A. G., Bradtke, S. J., and Singh, S. P · 1995
Earlier work this paper cites.
Optimal adaptive policies for Markov decision processes
Burnetas, A. N. and Katehakis, M. N · 1997
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Strens, M · 2000
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
Kernel-based reinforcement learning
Ormoneit, D. and Sen, Ś · 2002
Earlier work this paper cites.
Exploration in metric state spaces
Kakade, S., Kearns, M. J., and Langford, J · 2003
Earlier work this paper cites.
Self-normalized processes: Limit theory and Statistical Applications
Peña, V. H., Lai, T. L., and Shao, Q.-M · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Abbasi-Yadkori, Y. and Szepesvári, C · 2011
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Kernel-based reinforcement learning on representative states
Kveton, B. and Theocharous, G · 2012
Earlier work this paper cites.
Online regret bounds for undiscounted continuous reinforcement learning
Ortner, R. and Ryabko, D · 2012
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Boucheron, S., Lugosi, G., and Massart, P · 2013
Earlier work this paper cites.
The sample-complexity of general reinforcement learning
Lattimore, T., Hutter, M., Sunehag, P., et al · 2013
Earlier work this paper cites.