Fetching the paper…
Reading the bibliography…
We provide improved gap-dependent regret bounds for reinforcement learning in finite episodic Markov decision processes.
On tail probabilities for martingales
David A Freedman · 1975
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Efficient reinforcement learning
Claude-Nicolas Fiechter · 1994
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin Puterman · 1994
Earlier work this paper cites.
Asymptotically efficient adaptive choice of control laws incontrolled markov chains
Todd L Graves and Tze Leung Lai · 1997
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Kakade · 2003
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2007
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible MDPs
Ambuj Tewari and Peter L Bartlett · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2009
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Optimism in reinforcement learning and Kullback-Leibler divergence
Sarah Filippi, Olivier Cappé, and Aurélien Garivier · 2010
Earlier work this paper cites.
On the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2012
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Online learning in episodic markovian decision processes by relative entropy policy search
Alexander Zimin and Gergely Neu · 2013
Cited alongside, same era.
Pac reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Minimal exploration in structured stochastic bandits
Richard Combes, Stefan Magureanu, and Alexandre Proutiere · 2017
Cited alongside, same era.
Unifying PAC and regret: Uniform pac bounds for episodic reinforcement learning
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2019
Later among the works it cites.
Explore first, exploit next: The true shape of regret in bandit problems
Aurélien Garivier, Pierre Ménard, and Gilles Stoltz · 2019
Later among the works it cites.
Corruption robust exploration in episodic reinforcement learning
Thodoris Lykouris, Max Simchowitz, Aleksandrs Slivkins, and Wen Sun · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular MDPs
Max Simchowitz and Kevin Jamieson · 2019
Later among the works it cites.
A. Zanette and E. Brunskill · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
On oracle-efficient PAC reinforcement learning with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2018
Cited alongside, same era.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Nan Jiang and Alekh Agarwal · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Exploration in structured reinforcement learning
Jungseul Ok, Alexandre Proutiere, and Damianos Tranos · 2018
Cited alongside, same era.
Strategic Exploration in Reinforcement Learning - New Algorithms and Learning Guarantees
Christoph Dann · 2019
Cited alongside, same era.
Later among the works it cites.
Simon S Du, Jason D Lee, Gaurav Mahajan, and Ruosong Wang · 2020
Later among the works it cites.
Dylan J Foster, Alexander Rakhlin, David Simchi-Levi, and Yunzong Xu · 2020
Later among the works it cites.
Logarithmic regret for reinforcement learning with linear function approximation
Jiafan He, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Simultaneously learning stochastic and adversarial episodic MDPs with known transition
Tiancheng Jin and Haipeng Luo · 2020
Later among the works it cites.
Q Q -learning with logarithmic regret
Kunhe Yang, Lin F Yang, and Simon S Du · 2020
Later among the works it cites.
Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Haike Xu, Tengyu Ma, and Simon S Du · 2021
Closest in time.