Fetching the paper…
Reading the bibliography…
We propose a practical non-episodic PSRL algorithm that unlike recent state-of-the-art PSRL algorithms uses a deterministic, model-independent episode switching schedule.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Dynamic programming and optimal control
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L. Strehl and Michael L. Littman · 2005
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Peter L Bartlett and Ambuj Tewari · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Cited alongside, same era.
(More) efficient reinforcement learning via posterior sampling
Ian Osband, Dan Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Cited alongside, same era.
Bayesian optimal control of smoothly parameterized systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2015
Cited alongside, same era.
Thompson sampling for learning parameterized Markov decision processes
Aditya Gopalan and Shie Mannor · 2015
Later among the works it cites.
Posterior sampling for reinforcement learning without episodes
Ian Osband and Benjamin Van Roy · 2016
Later among the works it cites.
Posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Closest in time.
Learning-based control of unknown linear systems with thompson sampling
Yi Ouyang, Mukul Gagrani, and Rahul Jain · 2017
Closest in time.
Learning unknown Markov decision processes: A thompson sampling approach
Yi Ouyang, Mukul Gagrani, Ashutosh Nayyar, and Rahul Jain · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Georgios Theocharous, Nikos Vlassis, and Zheng Wen · 2017
Closest in time.