Fetching the paper…
Reading the bibliography…
We present an algorithm based on the \emph{Optimism in the Face of Uncertainty} (OFU) principle which is able to learn Reinforcement Learning (RL) modeled by Markov decision process (MDP) with finite state-action space efficiently.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
M L Puterman · 1994
Earlier work this paper cites.
Optimal Adaptive Policies for Markov Decision Processes
A. N. Burnetas and M. N. Katehakis · 1997
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Online regret bounds for markov decision processes with deterministic transitions
Ronald Ortner · 2008
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L Bartlett and Ambuj Tewari · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
A finite-time analysis of multi-armed bandits problems with kullback-leibler divergences
Odalric-Ambrym Maillard, Rémi Munos, and Gilles Stoltz · 2011
Earlier work this paper cites.
Pac bounds for discounted mdps
Tor Lattimore and Marcus Hutter · 2012
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Bayesian optimal control of smoothly parameterized systems
Yasin Abbasi-Yadkori · 2015
Cited alongside, same era.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2016
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning, worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Cited alongside, same era.
Posterior sampling for large scale reinforcement learning
Georgios Theocharous, Zheng Wen, Yasin Abbasi-Yadkori, and Nikos Vlassis · 2017
Later among the works it cites.
Andrea Zanette and Emma Brunskill · 2017
Later among the works it cites.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Nan Jiang and Alekh Agarwal · 2018
Later among the works it cites.
Variance reduction methods for sublinear reinforcement learning
Sham Kakade, Mengdi Wang, and Lin F Yang · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Learning unknown markov decision processes: A thompson sampling approach
Yi Ouyang, Mukul Gagrani, Ashutosh Nayyar, and Rahul Jain · 2017
Cited alongside, same era.
Near optimal exploration-exploitation in non-communicating markov decision processes
Ronan Fruit, Matteo Pirotta, and Alessandro Lazaric
Cited in the paper.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Ronan Fruit, Matteo Pirotta, Alessandro Lazaric, and Ronald Ortner
Cited in the paper.
Variance-aware regret bounds for undiscounted reinforcement learning in mdps
Mohammad Sadegh Talebi and Odalric-Ambrym Maillard · 2018
Later among the works it cites.
Improved analysis of ucrl2b
Ronan Fruit, Matteo Pirotta, and Alessandro Lazaric · 2019
Closest in time.