Fetching the paper…
Reading the bibliography…
Optimistic algorithms have been extensively studied for regret minimization in episodic tabular MDPs, both from a minimax and an instance-dependent view.
Efficient reinforcement learning
Claude-Nicolas Fiechter · 1994
Earlier work this paper cites.
Markov Decision Processes. Discrete Stochastic. Dynamic Programming
M.L. Puterman · 1994
Earlier work this paper cites.
lil’UCB: an Optimal Exploration Algorithm for Multi-Armed Bandits
K. Jamieson, M. Malloy, R. Nowak, and S. Bubeck · 2014
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Earlier work this paper cites.
The end of optimism? an asymptotic analysis of finite-armed linear bandits
Tor Lattimore and Csaba Szepesvari · 2017
Earlier work this paper cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sébastien Bubeck, and Michael I. Jordan · 2018
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin G. Jamieson · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Cited alongside, same era.
Planning in markov decision processes with gap-dependent sample complexity
Anders Jonsson, Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Edouard Leurent, and Michal Valko · 2020
Cited alongside, same era.
A unifying view of optimism in episodic reinforcement learning
Gergely Neu and Ciara Pike-Burke · 2020
Cited alongside, same era.
Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
Omar Darwiche Domingues, Pierre Ménard, Emilie Kaufmann, and Michal Valko · 2021
Later among the works it cites.
Adaptive reward-free exploration
Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Fast active learning for pure exploration in reinforcement learning
Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Beyond no regret: Instance-dependent PAC reinforcement learning
Andrew Wagenmaker, Max Simchowitz, and Kevin G. Jamieson · 2021
Later among the works it cites.
Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Haike Xu, Tengyu Ma, and Simon S Du · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christoph Dann, Teodor V. Marinov, Mehryar Mohri, and Julian Zimmert · 2021
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos
Cited in the paper.
Near instance-optimal PAC reinforcement learning for deterministic mdps
Andrea Tirinzoni, Aymen Al Marjani, and Emilie Kaufmann · 2022
Closest in time.