Fetching the paper…
Reading the bibliography…
We consider reinforcement learning in parameterized Markov Decision Processes (MDPs), where the parameterization may induce correlation across transition probabilities or rewards.
William R Thompson · 1933
Earlier work this paper cites.
Boundary crossing probabilities for the Wiener process and sample sums
Herbert Robbins and David Siegmund · 1970
Earlier work this paper cites.
Optimal control of a queueing system with two heterogeneous servers
Woei Lin and P.R. Kumar · 1984
Earlier work this paper cites.
Stochastic systems: estimation, identification, and adaptive control
P.R. Kumar and P.P. Varaiya · 1986
Earlier work this paper cites.
Asymptotically efficient adaptive allocation schemes for controlled Markov chains: finite parameter space
R. Agrawal, D. Teneketzis, and V. Anantharam · 1989
Earlier work this paper cites.
Probability and Random Processes
Geoffrey Grimmett and David Stirzaker · 1992
Earlier work this paper cites.
A simple proof of the optimality of a threshold policy in a two-server queueing system
Ger Koole · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
L.P. Kaelbling, M.L. Littman, and Andrew Moore · 1996
Earlier work this paper cites.
Optimal adaptive policies for Markov decision processes
Apostolos N. Burnetas and Michael N. Katehakis · 1997
Earlier work this paper cites.
Information-theoretic characterization of Bayes Performance and the Choice of Priors in Parametric and Nonparametric Problems
Andrew R. Barron · 1998
Earlier work this paper cites.
Model based Bayesian exploration
Richard Dearden, Nir Friedman, and David Andre · 1999
Earlier work this paper cites.
Posterior consistency of Dirichlet mixtures in density estimation
S. Ghosal, J. K. Ghosh, and R. V. Ramamoorthi · 1999
Earlier work this paper cites.
Convergence rates of posterior distributions
Subhashis Ghosal, Jayanta K. Ghosh, and Aad W. van der Vaart · 2000
Cited alongside, same era.
Rates of convergence of posterior distributions
Xiaotong Shen and Larry Wasserman · 2001
Cited alongside, same era.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2003
Cited alongside, same era.
Concentration inequalities
Stéphane Boucheron, Gábor Lugosi, and Olivier Bousquet · 2004
Cited alongside, same era.
Markov Chains and Mixing Times
David A. Levin, Yuval Peres, and Elizabeth L. Wilmer · 2006
Cited alongside, same era.
Pseudo-maximization and self-normalized processes
Victor H. de la Peña, Michael J. Klass, and Tze Leung Lai · 2007
Cited alongside, same era.
Near-optimal Regret Bounds for Reinforcement Learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Later among the works it cites.
A Minimum Relative Entropy Principle for Learning and Acting
P A Ortega and D A Braun · 2010
Later among the works it cites.
Analysis of Thompson sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal · 2012
Later among the works it cites.
Thompson Sampling: An Asymptotically Optimal Finite-time Analysis
Emilie Kaufmann, Nathaniel Korda, and Rémi Munos · 2012
Later among the works it cites.
Thompson Sampling for Contextual Bandits with Linear Payoffs
Shipra Agrawal and Navin Goyal · 2013
Later among the works it cites.
Thompson Sampling for 1-Dimensional Exponential Family Bandits
Nathaniel Korda, Emilie Kaufmann, and Remi Munos · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Remarks on consistency of posterior distributions , volume 3 of Collections
Taeryon Choi and R. V. Ramamoorthi · 2008
Cited alongside, same era.
Efficient reinforcement learning in parameterized models: Discrete parameter case
Kirill Dyagilev, Shie Mannor, and Nahum Shimkin · 2008
Cited alongside, same era.
An analysis of reinforcement learning with function approximation
Francisco S Melo, Sean P Meyn, and M Isabel Ribeiro · 2008
Cited alongside, same era.
Optimistic linear programming gives logarithmic regret for irreducible MDPs
Ambuj Tewari and Peter L. Bartlett · 2008
Cited alongside, same era.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
P.L. Bartlett and A. Tewari · 2009
Cited alongside, same era.
Regret bounds and minimax policies under partial monitoring
Jean-Yves Audibert and Sébastien Bubeck · 2010
Cited alongside, same era.
Computing the Stationary Distribution Locally
Christina E Lee, Asuman Ozdaglar, and Devavrat Shah · 2013
Later among the works it cites.
(More) Efficient Reinforcement Learning via Posterior Sampling
Ian Osband, Dan Russo, and Benjamin Van Roy · 2013
Later among the works it cites.
Eluder Dimension and the Sample Complexity of Optimistic Exploration
Dan Russo and Benjamin Van Roy · 2013
Later among the works it cites.
Thompson Sampling for Complex Online Problems
Aditya Gopalan, Shie Mannor, and Yishay Mansour · 2014
Closest in time.
Model-based Reinforcement Learning and the Eluder dimension
Ian Osband and Benjamin V. Roy · 2014
Closest in time.