Fetching the paper…
Reading the bibliography…
Bayesian model-based reinforcement learning is a formally elegant approach to learning optimal behaviour under model uncertainty, trading off exploration and exploitation in an ideal way.
On adaptive control processes
R. Bellman and R. Kalaba · 1959
Earlier work this paper cites.
Dual control theory
AA Feldbaum · 1960
Earlier work this paper cites.
Bayesian decision problems and Markov chains
J.J. Martin · 1967
Earlier work this paper cites.
Multi-armed bandit allocation indices
J.C. Gittins, R. Weber, and K.D. Glazebrook · 1989
Earlier work this paper cites.
Bayesian Q-learning
R. Dearden, N. Friedman, and S. Russell · 1998
Earlier work this paper cites.
Efficient Bayesian parameter estimation in large discrete domains
N. Friedman and Y. Singer · 1999
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large Markov decision processes
M. Kearns, Y. Mansour, and A.Y. Ng · 1999
Earlier work this paper cites.
Exploration of multi-state environments: Local measures and back-propagation of uncertainty
N. Meuleau and P. Bourgine · 1999
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
M. Strens · 2000
Earlier work this paper cites.
Optimal Learning: Computational Procedures For Bayes-Adaptive Markov Decision Processes
M.O.G. Duff · 2002
Cited alongside, same era.
Bayesian sparse sampling for on-line reward optimization
T. Wang, D. Lizotte, M. Bowling, and D. Schuurmans · 2005
Cited alongside, same era.
Bandit based Monte-Carlo planning
L. Kocsis and C. Szepesvári · 2006
Cited alongside, same era.
Bayesian exploration in Markov decision processes
P.S. Castro · 2007
Cited alongside, same era.
Bandit algorithms for tree search
P.A. Coquelin and R. Munos · 2007
Cited alongside, same era.
Combining online and offline knowledge in UCT
S. Gelly and D. Silver · 2007
Cited alongside, same era.
Smarter sampling in model-based Bayesian reinforcement learning
P. Castro and D. Precup · 2010
Later among the works it cites.
Monte-Carlo planning in large POMDPs
D. Silver and J. Veness · 2010
Later among the works it cites.
Variance-based rewards for approximate Bayesian reinforcement learning
J. Sorg, S. Singh, and R.L. Lewis · 2010
Later among the works it cites.
Algorithms for reinforcement learning
C. Szepesvári · 2010
Later among the works it cites.
Integrating sample-based planning and model-based reinforcement learning
T.J. Walsh, S. Goschin, and M.L. Littman · 2010
Later among the works it cites.
Approaching Bayes-optimality using Monte-Carlo tree search
J. Asmuth and M. Littman · 2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Asmuth, L. Li, M.L. Littman, A. Nouri, and D. Wingate · 2009
Cited alongside, same era.
Pure exploration in multi-armed bandits problems
S. Bubeck, R. Munos, and G. Stoltz · 2009
Cited alongside, same era.
Near-Bayesian exploration in polynomial time
J.Z. Kolter and A.Y. Ng · 2009
Cited alongside, same era.
The grand challenge of computer Go: Monte Carlo tree search and extensions
S. Gelly, L. Kocsis, M. Schoenauer, M. Sebag, D. Silver, C. Szepesvári, and O. Teytaud · 2012
Closest in time.
Scalable and efficient Bayes-adaptive reinforcement learning based on Monte-Carlo tree search
A. Guez, D. Silver, and P. Dayan · 2013
Closest in time.