Fetching the paper…
Reading the bibliography…
This paper establishes that optimistic algorithms attain gap-dependent and non-asymptotic logarithmic regret for episodic MDPs.
Optimal adaptive policies for markov decision processes
Apostolos N Burnetas and Michael N Katehakis · 1997
Earlier work this paper cites.
Reinforcement learning in large or unknown MDPs
Ambuj Tewari · 2007
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible mdps
Ambuj Tewari and Peter L Bartlett · 2008
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L Bartlett and Ambuj Tewari · 2009
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Ian Osband and Benjamin Van Roy · 2016
Cited alongside, same era.
Best-of-k-bandits
Max Simchowitz, Kevin Jamieson, and Benjamin Recht · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2018
Explore first, exploit next: The true shape of regret in bandit problems
Aurélien Garivier, Pierre Ménard, and Gilles Stoltz · 2018
Later among the works it cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Exploration in structured reinforcement learning
Jungseul Ok, Alexandre Proutiere, and Damianos Tranos · 2018
Later among the works it cites.
Efficient model-free reinforcement learning in metric spaces
Zhao Song and Wen Sun · 2019
Closest in time.
Andrea Zanette and Emma Brunskill · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.