Fetching the paper…
Reading the bibliography…
We study the reinforcement learning problem for discounted Markov Decision Processes (MDPs) under the tabular setting.
Q-learning with ucb exploration is sample efficient for infinite-horizon mdp
Dong, K · 1901
Earlier work this paper cites.
Zanette, A · 1901
Earlier work this paper cites.
Model-based reinforcement learning with a generative model is minimax optimal
Agarwal, A · 1906
Earlier work this paper cites.
Variance-reduced q q -learning is minimax optimal
Wainwright, M. J · 1906
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Kearns, M. J · 1999
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M · 2003
Earlier work this paper cites.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zhang, Z · 2004
Earlier work this paper cites.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2006
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N · 2006
Earlier work this paper cites.
On optimism in model-based reinforcement learning
Pacchiano, A · 2006
Earlier work this paper cites.
Pac model-free reinforcement learning
Strehl, A. L · 2006
Cited alongside, same era.
Model-free reinforcement learning: from clipped pseudo-regret to sample complexity
Zhang, Z · 2006
Cited alongside, same era.
A unifying view of optimism in episodic reinforcement learning
Neu, G · 2007
Cited alongside, same era.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L · 2008
Cited alongside, same era.
Empirical bernstein bounds and sample variance penalization
Maurer, A · 2009
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Osband, I · 2016
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Azar, M. G · 2017
Later among the works it cites.
Why is posterior sampling better than optimism for reinforcement learning?
Osband, I · 2017
Later among the works it cites.
Wang, M · 2017
Later among the works it cites.
Is q-learning provably efficient?
Jin, C · 2018
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Dann, C · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jaksch, T · 2010
Cited alongside, same era.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Szita, I · 2010
Cited alongside, same era.
Pac bounds for discounted mdps
Lattimore, T · 2012
Cited alongside, same era.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G · 2013
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C · 2015
Cited alongside, same era.
Sidford, A
Cited in the paper.
Variance reduced value iteration and faster algorithms for solving markov decision processes
Sidford, A
Cited in the paper.
Later among the works it cites.
Worst-case regret bounds for exploration via randomized value functions
Russo, D · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M · 2019
Later among the works it cites.
Regret bounds for discounted mdps
Liu, S · 2020
Closest in time.
Q-learning with logarithmic regret 1576–1584
Yang, K · 2021
Closest in time.