Fetching the paper…
Reading the bibliography…
Computational results demonstrate that posterior sampling for reinforcement learning (PSRL) dramatically outperforms algorithms driven by optimism, such as UCRL2.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W.R · 1933
Earlier work this paper cites.
On adaptive control processes
Bellman, Richard and Kalaba, Robert · 1959
Earlier work this paper cites.
Rules for ordering uncertain prospects
Hadar, Josef and Russell, William R · 1969
Earlier work this paper cites.
Sub-gaussian random variables
Buldygin, Valerii V and Kozachenko, Yu V · 1980
Earlier work this paper cites.
Optimal adaptive policies for Markov decision processes
Burnetas, Apostolos N and Katehakis, Michael N · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard and Barto, Andrew · 1998
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Strens, Malcolm J. A · 2000
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, Ronen I. and Tennenholtz, Moshe · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, Michael J. and Singh, Satinder P · 2002
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Kakade, Sham · 2003
Earlier work this paper cites.
A theoretical analysis of model-based interval estimation
Strehl, Alexander L and Littman, Michael L · 2005
Earlier work this paper cites.
PAC model-free reinforcement learning
Strehl, Alexander L., Li, Lihong, Wiewiora, Eric, Langford, John, and Littman, Michael L · 2006
Cited alongside, same era.
Stochastic linear optimization under bandit feedback
Dani, Varsha, Hayes, Thomas P., and Kakade, Sham M · 2008
Cited alongside, same era.
A Bayesian sampling approach to exploration in reinforcement learning
Asmuth, John, Li, Lihong, Littman, Michael L, Nouri, Ali, and Wingate, David · 2009
Cited alongside, same era.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Bartlett, Peter L. and Tewari, Ambuj · 2009
Cited alongside, same era.
Near-Bayesian exploration in polynomial time
Kolter, J Zico and Ng, Andrew Y · 2009
Cited alongside, same era.
Optimism in reinforcement learning and kullback-leibler divergence
Filippi, Sarah, Cappé, Olivier, and Garivier, Aurélien · 2010
An optimistic posterior sampling strategy for Bayesian reinforcement learning
Fonteneau, Raphaël, Korda, Nathan, and Munos, Rémi · 2013
Later among the works it cites.
(More) efficient reinforcement learning via posterior sampling
Osband, Ian, Russo, Daniel, and Van Roy, Benjamin · 2013
Later among the works it cites.
Thompson sampling for learning parameterized Markov decision processes
Gopalan, Aditya and Mannor, Shie · 2014
Later among the works it cites.
From bandits to monte-carlo tree search: The optimistic principle applied to optimization and planning
Munos, Rémi · 2014
Later among the works it cites.
Generalization and exploration via randomized value functions
Osband, Ian, Van Roy, Benjamin, and Wen, Zheng · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, Thomas, Ortner, Ronald, and Auer, Peter · 2010
Cited alongside, same era.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Szita, István and Szepesvári, Csaba · 2010
Cited alongside, same era.
Near-optimal brl using optimistic local transitions
Araya, Mauricio, Buffet, Olivier, and Thomas, Vincent · 2012
Cited alongside, same era.
PAC bounds for discounted MDPs
Lattimore, Tor and Hutter, Marcus · 2012
Cited alongside, same era.
Bayesian reinforcement learning
Vlassis, Nikos, Ghavamzadeh, Mohammad, Mannor, Shie, and Poupart, Pascal · 2012
Cited alongside, same era.
Model-based reinforcement learning and the eluder dimension
Osband, Ian and Van Roy, Benjamin
Cited in the paper.
Learning to optimize via posterior sampling
Russo, Daniel and Van Roy, Benjamin · 2014
Later among the works it cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, Christoph and Brunskill, Emma · 2015
Later among the works it cites.
Deep Exploration via Randomized Value Functions
Osband, Ian · 2016
Closest in time.
On lower bounds for regret in reinforcement learning
Osband, Ian and Van Roy, Benjamin · 2016
Closest in time.
Gaussian-dirichlet posterior dominance in sequential learning
Osband, Ian and Van Roy, Benjamin · 2017
Closest in time.