Fetching the paper…
Reading the bibliography…
Episodic reinforcement learning and contextual bandits are two widely studied sequential decision-making problems.
On tail probabilities for martingales
David A Freedman · 1975
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham M Kakade · 2003
Earlier work this paper cites.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zihan Zhang, Yuan Zhou, and Xiangyang Ji · 2004
Earlier work this paper cites.
Pac model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Model-free reinforcement learning: from clipped pseudo-regret to sample complexity
Zihan Zhang, Yuan Zhou, and Xiangyang Ji · 2006
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Regal: a regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L Bartlett and Ambuj Tewari · 2009
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
J Zico Kolter and Andrew Y Ng · 2009
Earlier work this paper cites.
Empirical Bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
István Szita and Csaba Szepesvári · 2010
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sebastien Bubeck and Nicolo Cesa-Bianchi · 2012
Cited alongside, same era.
Pac bounds for discounted mdps
Tor Lattimore and Marcus Hutter · 2012
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Ian Osband and Benjamin Van Roy · 2016
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
On the optimality of sparse model-based planning for Markov decision processes
Alekh Agarwal, Sham Kakade, and Lin F Yang · 2019
Later among the works it cites.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2019
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2019
Later among the works it cites.
Q-learning with ucb exploration is sample efficient for infinite-horizon mdp
Kefan Dong, Yuanhao Wang, Xiaoyu Chen, and Liwei Wang · 2019
Later among the works it cites.
Worst-case regret bounds for exploration via randomized value functions
Daniel Russo · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2017
Cited alongside, same era.
Near optimal exploration-exploitation in non-communicating markov decision processes
Ronan Fruit, Matteo Pirotta, and Alessandro Lazaric · 2018
Cited alongside, same era.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Nan Jiang and Alekh Agarwal · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin G Jamieson · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zihan Zhang and Xiangyang Ji · 2019
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Closest in time.
A unifying view of optimism in episodic reinforcement learning
Gergely Neu and Ciara Pike-Burke · 2020
Closest in time.
On optimism in model-based reinforcement learning
Aldo Pacchiano, Philip Ball, Jack Parker-Holder, Krzysztof Choromanski, and Stephen Roberts · 2020
Closest in time.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?
Ruosong Wang, Simon S Du, Lin F Yang, and Sham M Kakade · 2020
Closest in time.
Q Q -learning with logarithmic regret
Kunhe Yang, Lin F Yang, and Simon S Du · 2020
Closest in time.