Fetching the paper…
Reading the bibliography…
We consider model-based reinforcement learning in finite Markov De- cision Processes (MDPs), focussing on so-called optimistic strategies.
Asymptotically efficient adaptive allocation rules
T. Lai and H. Robbins · 1985
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. Puterman · 1994
Earlier work this paper cites.
Optimal adaptive policies for Markov decision processes
A. Burnetas and M. Katehakis · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
R. Sutton and A. Barto · 1998
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
R. Brafman and M. Tennenholtz · 2003
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition matrices
A. Nilim and L. El Ghaoui · 2005
Cited alongside, same era.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
E. Even-Dar, S. Mannor, and Y. Mansour · 2006
Cited alongside, same era.
Logarithmic online regret bounds for undiscounted reinforcement learning
P. Auer and R. Ortner · 2007
Cited alongside, same era.
On upper-confidence bound policies for non-stationary bandit problems
A. Garivier and E. Moulines · 2008
Cited alongside, same era.
An analysis of model-based interval estimation for Markov decision processes
A. Strehl and M. Littman · 2008
Cited alongside, same era.
Optimistic linear programming gives logarithmic regret for irreducible MDPs
A. Tewari and P. Bartlett · 2008
Later among the works it cites.
Near-optimal regret bounds for reinforcement learning
P. Auer, T. Jaksch, and R. Ortner · 2009
Later among the works it cites.
REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs
P. Bartlett and A. Tewari · 2009
Later among the works it cites.
Context tree selection: A unifying view
A. Garivier and F. Leonardi · 2010
Closest in time.
Near-optimal Regret Bounds for Reinforcement Learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…