Fetching the paper…
Reading the bibliography…
We study upper and lower bounds on the sample-complexity of learning near-optimal behaviour in finite-state discounted Markov Decision Processes (MDPs).
On a modification of Chebyshev’s inequality and of the error formula of Laplace
S. Bernstein · 1924
Earlier work this paper cites.
The variance of discounted Markov decision processes
M. Sobel · 1982
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. Lai and H. Robbins · 1985
Earlier work this paper cites.
Sample mean based index policies with O(log n) regret for the multi-armed bandit problem
R. Agrawal · 1995
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
On The Sample Complexity Of Reinforcement Learning
S. Kakade · 2003
Cited alongside, same era.
The sample complexity of exploration in the multi-armed bandit problem
S. Mannor and J. Tsitsiklis · 2004
Cited alongside, same era.
Model-based reinforcement learning with nearly tight exploration complexity bounds
I. Szita and C. Szepesvári · 2004
Cited alongside, same era.
A theoretical analysis of model-based interval estimation
A. Strehl and M. Littman · 2005
Cited alongside, same era.
PAC model-free reinforcement learning
A. Strehl, L. Li, E. Wiewiorac, J. Langford, and M. Littman · 2006
Cited alongside, same era.
Logarithmic online regret bounds for undiscounted reinforcement learning
P. Auer and R. Ortner · 2007
Later among the works it cites.
An analysis of model-based interval estimation for Markov decision processes
A. Strehl and M. Littman · 2008
Later among the works it cites.
Reinforcement learning in finite MDPs: PAC analysis
A. Strehl, L. Li, and M. Littman · 2009
Later among the works it cites.
Near-optimal regret bounds for reinforcement learning
P. Auer, T. Jaksch, and R. Ortner · 2010
Later among the works it cites.
Upper confidence reinforcement learning
P. Auer · 2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…