Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has traditionally been understood from an episodic perspective; the concept of non-episodic RL, where there is no restart and therefore no reliable recovery, remains elusive.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 1995
Earlier work this paper cites.
Average reward reinforcement learning: Foundations, algorithms, and empirical results
S. Mahadevan · 1996
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade et al · 2003
Earlier work this paper cites.
Double q-learning
H. V. Hasselt · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
I. Szita and C. Szepesvári · 2010
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
P. L. Bartlett and A. Tewari · 2012
Cited alongside, same era.
Pac bounds for discounted mdps
T. Lattimore and M. Hutter · 2012
Cited alongside, same era.
Deep reinforcement learning with double q-learning
H. V. Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
M. G. Azar, I. Osband, and R. Munos · 2017
Cited alongside, same era.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
R. Fruit, M. Pirotta, A. Lazaric, and R. Ortner · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. V. Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Is q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Later among the works it cites.
Strategic Exploration in Reinforcement Learning-New Algorithms and Learning Guarantees
C. Dann · 2019
Later among the works it cites.
Q-learning with ucb exploration is sample efficient for infinite-horizon mdp
K. Dong, Y. Wang, X. Chen, and L. Wang · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
A. Zanette and E. Brunskill · 2019
Later among the works it cites.
Nearly minimax optimal reinforcement learning for discounted mdps
J. He, D. Zhou, and Q. Gu · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Regret bounds for reinforcement learning via markov chain concentration
R. Ortner · 2020
Closest in time.