Fetching the paper…
Reading the bibliography…
We introduce and analyse two algorithms for exploration-exploitation in discrete and continuous Markov Decision Processes (MDPs) based on exploration bonuses.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 1935
Earlier work this paper cites.
On tail probabilities for martingales
David A. Freedman · 1975
Earlier work this paper cites.
Probability Theory: Independence, Interchangeability, Martingales
Y.S. Chow and H. Teicher · 1988
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Dynamic programming and optimal control. Vol II
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Improved risk tail bounds for on-line algorithms
Nicolò Cesa-Bianchi and Claudio Gentile · 2005
Earlier work this paper cites.
Optimism in the face of uncertainty should be refutable
Ronald Ortner · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Peter L. Bartlett and Ambuj Tewari · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Online regret bounds for undiscounted continuous reinforcement learning
Ronald Ortner and Daniil Ryabko · 2012
Cited alongside, same era.
Probability Theory: A Comprehensive Course
A. Klenke and M. Loève · 2013
Cited alongside, same era.
Improved regret bounds for undiscounted continuous reinforcement learning
K. Lakshmanan, Ronald Ortner, and Daniil Ryabko · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos · 2016
Regret minimization in mdps with options without prior knowledge
Ronan Fruit, Matteo Pirotta, Alessandro Lazaric, and Emma Brunskill · 2017
Later among the works it cites.
Count-based exploration in feature space for reinforcement learning
Jarryd Martin, Suraj Narayanan Sasikumar, Tom Everitt, and Marcus Hutter · 2017
Later among the works it cites.
Count-based exploration with neural density models
Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Later among the works it cites.
#exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2017
Later among the works it cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sébastien Bubeck, and Michael I. Jordan · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Near optimal exploration-exploitation in non-communicating markov decision processes
Ronan Fruit, Matteo Pirotta, and Alessandro Lazaric
Cited in the paper.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Ronan Fruit, Matteo Pirotta, Alessandro Lazaric, and Ronald Ortner
Cited in the paper.
Closest in time.
Variance reduction methods for sublinear reinforcement learning
Sham Kakade, Mengdi Wang, and Lin F. Yang · 2018
Closest in time.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2018
Closest in time.
Variance-aware regret bounds for undiscounted reinforcement learning in mdps
Mohammad Sadegh Talebi and Odalric-Ambrym Maillard · 2018
Closest in time.