Fetching the paper…
Reading the bibliography…
We consider the problem of provably optimal exploration in reinforcement learning for finite horizon MDPs.
Theory of probability, 1927
Bernstein, S · 1927
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W.R · 1933
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, David A · 1975
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N · 1996
Earlier work this paper cites.
Optimal adaptive policies for markov decision processes
Burnetas, Apostolos N and Katehakis, Michael N · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard and Barto, Andrew · 1998
Earlier work this paper cites.
Influence and variance of a Markov chain : Application to adaptive discretizations in optimal control
Munos, R. and Moore, A · 1999
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Strens, Malcolm J. A · 2000
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, Ronen I. and Tennenholtz, Moshe · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, Michael J. and Singh, Satinder P · 2002
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Weissman, Tsachy, Ordentlich, Erik, Seroussi, Gadiel, Verdu, Sergio, and Weinberger, Marcelo J · 2003
Cited alongside, same era.
A theoretical analysis of model-based interval estimation
Strehl, Alexander L and Littman, Michael L · 2005
Cited alongside, same era.
Prediction, Learning, and Games
Cesa-Bianchi, N. and Lugosi, G · 2006
Cited alongside, same era.
PAC model-free reinforcement learning
Strehl, Alexander L., Li, Lihong, Wiewiora, Eric, Langford, John, and Littman, Michael L · 2006
Cited alongside, same era.
Dynamic Programming and Optimal Control , volume I
Bertsekas, D. P · 2007
Cited alongside, same era.
An analysis of model-based interval estimation for markov decision processes
Strehl, Alexander L and Littman, Michael L · 2008
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, Sébastien and Cesa-Bianchi, Nicolò · 2012
Later among the works it cites.
PAC bounds for discounted MDPs
Lattimore, Tor and Hutter, Marcus · 2012
Later among the works it cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Azar, Mohammad Gheshlaghi, Munos, Rémi, and Kappen, Hilbert J · 2013
Later among the works it cites.
Scalable and efficient bayes-adaptive reinforcement learning based on monte-carlo tree search
Guez, Arthur, Silver, David, and Dayan, Peter · 2013
Later among the works it cites.
(more) efficient reinforcement learning via posterior sampling
Osband, Ian, Russo, Dan, and Van Roy, Benjamin · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Bartlett, Peter L. and Tewari, Ambuj · 2009
Cited alongside, same era.
Empirical bernstein bounds and sample variance penalization
Maurer, Andreas and Pontil, Massimiliano · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Cited alongside, same era.
X-armed bandits
Bubeck, Sébastien, Munos, Rémi, Stoltz, Gilles, and Szepesvári, Csaba · 2011
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Osband, Ian and Van Roy, Benjamin
Cited in the paper.
Why is posterior sampling better than optimism for reinforcement learning
Osband, Ian and Van Roy, Benjamin
Cited in the paper.
From bandits to Monte-Carlo Tree Search: The optimistic principle applied to optimization and planning
Munos, Rémi · 2014
Later among the works it cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, Christoph and Brunskill, Emma · 2015
Later among the works it cites.
Posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, Shipra and Jia, Randy · 2017
Closest in time.
Ubev-a more practical algorithm for episodic rl with near-optimal pac and regret guarantees
Dann, Christoph, Lattimore, Tor, and Brunskill, Emma · 2017
Closest in time.