Fetching the paper…
Reading the bibliography…
Most known regret bounds for reinforcement learning are either episodic or assume an environment without traps.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 1935
Earlier work this paper cites.
On integrating apprentice learning and reinforcement learning
J. Clouse · 1997
Earlier work this paper cites.
Handbook of Markov Decision Processes
Eugene A. Feinberg and Adam Shwartz (eds.) · 2002
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L. Bartlett · 2009
Cited alongside, same era.
Competing with an infinite set of models in reinforcement learning
Phuong Nguyen, Odalric-Ambrym Maillard, Daniil Ryabko, and Ronald Ortner · 2013
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Benjamin Van Roy, and Daniel Russo · 2013
Cited alongside, same era.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Later among the works it cites.
A comprehensive survey on safe reinforcement learning
Javier García and Fernando Fernández · 2015
Later among the works it cites.
An information-theoretic analysis of thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…