Fetching the paper…
Reading the bibliography…
We address reinforcement learning problems with finite state and action spaces where the underlying MDP has some known structure that could be potentially exploited to minimize the exploration rates of suboptimal (state, action) pairs.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 1994
Earlier work this paper cites.
Optimal adaptive policies for Markov decision processes
Apostolos N. Burnetas and Michael N. Katehakis · 1997
Earlier work this paper cites.
Asymptotically efficient adaptive choice of control laws in controlled Markov chains
Todd L. Graves and Tze Leung Lai · 1997
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2007
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible MDPs
Ambuj Tewari and Peter L. Bartlett · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2009
Earlier work this paper cites.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Peter L. Bartlett and Ambuj Tewari · 2009
Earlier work this paper cites.
Optimism in reinforcement learning and Kullback-Leibler divergence
Sarah Filippi, Olivier Cappé, and Aurélien Garivier · 2010
Cited alongside, same era.
Online regret bounds for undiscounted continuous reinforcement learning
Ronald Ortner and Daniil Ryabko · 2012
Cited alongside, same era.
Unimodal bandits: Regret lower bounds and optimal algorithms
Richard Combes and Alexandre Proutiere · 2014
Cited alongside, same era.
Lipschitz bandits: Regret lower bounds and optimal algorithms
Stefan Magureanu, Richard Combes, and Alexandre Proutiere · 2014
Cited alongside, same era.
Model-based reinforcement learning and the Eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Cited alongside, same era.
Improved regret bounds for undiscounted continuous reinforcement learning
Kailasam Lakshmanan, Ronald Ortner, and Daniil Ryabko · 2015
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
On the complexity of best-arm identification in multi-armed bandit models
Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier · 2016
Later among the works it cites.
On lower bounds for regret in reinforcement learning
Ian Osband and Benjamin Van Roy · 2016
Later among the works it cites.
Posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Later among the works it cites.
Minimal exploration in structured stochastic bandits
Richard Combes, Stefan Magureanu, and Alexandre Proutiere · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Explore first, exploit next: The true shape of regret in bandit problems
Aurélien Garivier, Pierre Ménard, and Gilles Stoltz · 2018
Closest in time.