Fetching the paper…
Reading the bibliography…
This is a brief technical note to clarify some of the issues with applying the application of the algorithm posterior sampling for reinforcement learning (PSRL) in environments without fixed episodes.
A Bayesian framework for reinforcement learning
Malcolm J. A. Strens · 2000
Earlier work this paper cites.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Peter L. Bartlett and Ambuj Tewari · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Earlier work this paper cites.
(More) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Thompson sampling for learning parameterized Markov decision processes
Aditya Gopalan and Shie Mannor · 2014
Cited alongside, same era.
From bandits to monte-carlo tree search: The optimistic principle applied to optimization and planning
Rémi Munos · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy
Cited in the paper.
Near-optimal reinforcement learning in factored MDPs
Ian Osband and Benjamin Van Roy
Cited in the paper.
Bayesian optimal control of smoothly parameterized systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2015
Later among the works it cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Later among the works it cites.
Why is posterior sampling better than optimism for reinforcement learning
Ian Osband and Benjamin Van Roy · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…