Fetching the paper…
Reading the bibliography…
Most provably-efficient learning algorithms introduce optimism about poorly-understood states and actions to encourage exploration.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T.L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Stochastic systems: estimation, identification and adaptive control
P. R. Kumar and P. Varaiya · 1986
Earlier work this paper cites.
Optimal adaptive policies for markov decision processes
A. N. Burnetas and M. N. Katehakis · 1997
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
M. Strens · 2000
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade · 2003
Earlier work this paper cites.
Bayesian sparse sampling for on-line reward optimization
T. Wang, D. Lizotte, M. Bowling, and D. Schuurmans · 2005
Cited alongside, same era.
An analysis of model-based interval estimation for markov decision processes
A. L. Strehl and M. L. Littman · 2008
Cited alongside, same era.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
P. L. Bartlett and A. Tewari · 2009
Cited alongside, same era.
Near-Bayesian exploration in polynomial time
J. Z. Kolter and A. Y. Ng · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Cited alongside, same era.
A modern Bayesian look at the multi-armed bandit
S.L. Scott · 2010
Cited alongside, same era.
An empirical evaluation of Thompson sampling
O. Chapelle and L. Li · 2011
Later among the works it cites.
Approaching bayes-optimalilty using monte-carlo tree search
J. Asmuth and M. L. Littman · 2011
Later among the works it cites.
Further optimal regret bounds for Thompson sampling
S. Agrawal and N. Goyal · 2012
Later among the works it cites.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal · 2012
Later among the works it cites.
Thompson sampling: an asymptotically optimal finite time analysis
E. Kauffmann, N. Korda, and R. Munos · 2012
Later among the works it cites.
Efficient bayes-adaptive reinforcement learning using sample-based search
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimism in reinforcement learning based on kullback-leibler divergence
S. Filippi, O. Cappé, and A. Garivier · 2010
Cited alongside, same era.
A. Guez, D. Silver, and P. Dayan · 2012
Later among the works it cites.
Learning to optimize via posterior sampling
D. Russo and B. Van Roy · 2013
Closest in time.