Fetching the paper…
Reading the bibliography…
We derive sublinear regret bounds for undiscounted reinforcement learning in continuous state space.
Discrete-time Markov control processes
Onésimo Hernández-Lerma and Jean Bernard Lasserre · 1996
Earlier work this paper cites.
Tree based discretization for continuous state space reinforcement learning
William T. B. Uther and Manuela M. Veloso · 1998
Earlier work this paper cites.
Further topics on discrete-time Markov control processes
Onésimo Hernández-Lerma and Jean Bernard Lasserre · 1999
Earlier work this paper cites.
Finite-time analysis of the multi-armed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael J. Kearns and Satinder P. Singh · 2002
Earlier work this paper cites.
Exploration in metric state spaces
Sham Kakade, Michael J. Kearns, and John Langford · 2003
Earlier work this paper cites.
Nearly tight bounds for the continuum-armed bandit problem
Robert Kleinberg · 2005
Earlier work this paper cites.
Improved rates for the stochastic continuum-armed bandit problem
Peter Auer, Ronald Ortner, and Csaba Szepesvári · 2007
Cited alongside, same era.
Model-based exploration in continuous state spaces
Nicholas K. Jong and Peter Stone · 2007
Cited alongside, same era.
Efficient continuous-time reinforcement learning with adaptive state graphs
Gerhard Neumann, Michael Pfeiffer, and Wolfgang Maass · 2007
Cited alongside, same era.
Multi-armed bandits in metric spaces
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal · 2008
Cited alongside, same era.
Online linear regression and its application to model-based reinforcement learning
Alexander L. Strehl and Michael L. Littman · 2008
Cited alongside, same era.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Peter L. Bartlett and Ambuj Tewari · 2009
Adaptive-resolution reinforcement learning with polynomial exploration in deterministic domains
Andrey Bernstein and Nahum Shimkin · 2010
Later among the works it cites.
Online optimization of χ \chi -armed bandits
Sébastien Bubeck, Rémi Munos, Gilles Stoltz, and Csaba Szepesvári · 2010
Later among the works it cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Later among the works it cites.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Later among the works it cites.
Selecting the state-representation in reinforcement learning
Odalric-Ambrym Maillard, Rémi Munos, and Daniil Ryabko · 2012
Later among the works it cites.
Adaptive aggregation for reinforcement learning in average reward Markov decision processes
Ronald Ortner · 2012
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Provably efficient learning with typed parametric models
Emma Brunskill, Bethany R. Leffler, Lihong Li, Michael L. Littman, and Nicholas Roy · 2009
Cited alongside, same era.
Multi-resolution exploration in continuous spaces
Ali Nouri and Michael L. Littman · 2009
Cited alongside, same era.
Later among the works it cites.
Regret bounds for restless Markov bandits
Ronald Ortner, Daniil Ryabko, Peter Auer, and Rémi Munos · 2012
Later among the works it cites.
Optimal regret bounds for selecting the state representation in reinforcement learning
Odalric-Ambrym Maillard, Phuong Nguyen, Ronald Ortner, and Daniil Ryabko · 2013
Closest in time.