Fetching the paper…
Reading the bibliography…
This paper presents a new model-free algorithm for episodic finite-horizon Markov Decision Processes (MDP), Adaptive Multi-step Bootstrap (AMB), which enjoys a stronger gap-dependent regret bound.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
PAC model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2007
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible MDPs
Ambuj Tewari and Peter L Bartlett · 2008
Earlier work this paper cites.
Zihan Zhang, Xiangyang Ji, and Simon S Du · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Algorithms for reinforcement learning
Csaba Szepesvári · 2010
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sebastien Bubeck and Nicolo Cesa-Bianchi · 2012
Cited alongside, same era.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Corruption robust exploration in episodic reinforcement learning
Thodoris Lykouris, Max Simchowitz, Aleksandrs Slivkins, and Wen Sun · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular MDPs
Max Simchowitz and Kevin G Jamieson · 2019
Later among the works it cites.
Introduction to multi-armed bandits
Aleksandrs Slivkins · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
Logarithmic regret for reinforcement learning with linear function approximation
Jiafan He, Dongruo Zhou, and Quanquan Gu · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nan Jiang and Alekh Agarwal · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Exploration in structured reinforcement learning
Jungseul Ok, Alexandre Proutiere, and Damianos Tranos · 2018
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2019
Cited alongside, same era.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zihan Zhang, Yuan Zhou, and Xiangyang Ji
Cited in the paper.
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?
Ruosong Wang, Simon S Du, Lin F Yang, and Sham M Kakade · 2020
Later among the works it cites.
Q Q -learning with Logarithmic Regret
Kunhe Yang, Lin F Yang, and Simon S Du · 2020
Later among the works it cites.