Fetching the paper…
Reading the bibliography…
We consider the problem of learning an unknown Markov Decision Process (MDP) that is weakly communicating in the infinite horizon setting.
W. R. Thompson, “On the likelihood that one unknown probability exceeds another in view of the evidence of two samples,” Biometrika
1933
Earlier work this paper cites.
T. L. Lai and H. Robbins, “Asymptotically efficient adaptive allocation rules,” Advances in applied mathematics
1985
Earlier work this paper cites.
A. N. Burnetas and M. N. Katehakis, “Optimal adaptive policies for markov decision processes,” Mathematics of Operations Research
1997
Earlier work this paper cites.
M. Strens, “A bayesian framework for reinforcement learning,” in ICML
2000
Earlier work this paper cites.
M. Kearns and S. Singh, “Near-optimal reinforcement learning in polynomial time,” Machine Learning
2002
Earlier work this paper cites.
R. I. Brafman and M. Tennenholtz, “R-max-a general polynomial time algorithm for near-optimal reinforcement learning,” Journal of Machine Learning Research
2002
Earlier work this paper cites.
2008
Earlier work this paper cites.
P. L. Bartlett and A. Tewari, “Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps,” in UAI
2009
Earlier work this paper cites.
T. Jaksch, R. Ortner, and P. Auer, “Near-optimal regret bounds for reinforcement learning,” Journal of Machine Learning Research
2010
Cited alongside, same era.
S. Filippi, O. Cappé, and A. Garivier, “Optimism in reinforcement learning and kullback-leibler divergence,” in Allerton
2010
Cited alongside, same era.
S. L. Scott, “A modern bayesian look at the multi-armed bandit,” Applied Stochastic Models in Business and Industry
2010
Cited alongside, same era.
O. Chapelle and L. Li, “An empirical evaluation of thompson sampling,” in NIPS
2011
Cited alongside, same era.
Athena Scientific, Belmont, MA, 2012
D. P. Bertsekas, Dynamic programming and optimal control · 2012
Cited alongside, same era.
I. Osband, D. Russo, and B. Van Roy, “(More) efficient reinforcement learning via posterior sampling,” in NIPS
D. Russo and B. Van Roy, “Learning to optimize via posterior sampling,” Mathematics of Operations Research
2014
Later among the works it cites.
SIAM, 2015
P. R. Kumar and P. Varaiya, Stochastic systems: Estimation, identification, and adaptive control · 2015
Later among the works it cites.
C. Dann and E. Brunskill, “Sample complexity of episodic fixed-horizon reinforcement learning,” in NIPS
2015
Later among the works it cites.
A. Gopalan and S. Mannor, “Thompson sampling for learning parameterized markov decision processes,” in COLT
2015
Later among the works it cites.
Y. Abbasi-Yadkori and C. Szepesvári, “Bayesian optimal control of smoothly parameterized systems.,” in UAI
2015
Later among the works it cites.
I. Osband and B. Van Roy, “Why is posterior sampling better than optimism for reinforcement learning,” EWRL
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
R. Fonteneau, N. Korda, and R. Munos, “An optimistic posterior sampling strategy for bayesian reinforcement learning,” in BayesOpt2013
2013
Cited alongside, same era.
2016
Later among the works it cites.
2016
Later among the works it cites.