Fetching the paper…
Reading the bibliography…
The principle of optimism in the face of uncertainty underpins many theoretically successful reinforcement learning algorithms.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Les problèmes de décisions séquentielles
G. de Ghellinck · 1960
Earlier work this paper cites.
Dynamic Programming and Markov Processes
R. A. Howard · 1960
Earlier work this paper cites.
Linear programming and sequential decisions
A. S. Manne · 1960
Earlier work this paper cites.
On linear programming in a Markov decision problem
E. V. Denardo · 1970
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
Sample mean based index policies with O ( l o g n ) {O}(logn) regret for the multi-armed bandit problem
R. Agrawal · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. J. Bradtke and A. G. Barto · 1996
Earlier work this paper cites.
Optimal adaptive policies for sequential allocation problems
A. Burnetas and M. Katehakis · 1996
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
R-MAX - a general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. Kakade · 2003
Earlier work this paper cites.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Earlier work this paper cites.
Learning near optimal policies with low inherent bellman error
A. Zanette, A. Lazaric, M. Kochenderfer, and E. Brunskill · 2003
Earlier work this paper cites.
Convex optimization
S. Boyd, S. P. Boyd, and L. Vandenberghe · 2004
Earlier work this paper cites.
On divergences and informations in statistics and information theory
F. Liese and I. Vajda · 2006
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
P. Auer and R. Ortner · 2007
Cited alongside, same era.
Dynamic Programming and Optimal Control , volume 2
D. P. Bertsekas · 2007
Cited alongside, same era.
Stochastic linear optimization under bandit feedback
V. Dani, T. P. Hayes, and S. M. Kakade · 2008
Cited alongside, same era.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
R. Parr, L. Li, G. Taylor, C. Painter-Wakefield, and M. L. Littman · 2008
Cited alongside, same era.
An analysis of model-based interval estimation for Markov decision processes
A. L. Strehl and M. L. Littman · 2008
Cited alongside, same era.
Optimistic linear programming gives logarithmic regret for irreducible MDPs
A. Tewari and P. L. Bartlett · 2008
On lower bounds for regret in reinforcement learning
I. Osband and B. Van Roy · 2016
Later among the works it cites.
Policy error bounds for model-based reinforcement learning with factored linear models
B. Á. Pires and C · 2016
Later among the works it cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
S. Agrawal and R. Jia · 2017
Later among the works it cites.
Minimax regret bounds for reinforcement learning
M. G. Azar, I. Osband, and R. Munos · 2017
Later among the works it cites.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
C. Dann, T. Lattimore, and E. Brunskill · 2017
Later among the works it cites.
Is Q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
REGAL: a regularization based algorithm for reinforcement learning in weakly communicating MDPs
P. L. Bartlett and A. Tewari · 2009
Cited alongside, same era.
Empirical bernstein bounds and sample variance penalization
A. Maurer and M. Pontil · 2009
Cited alongside, same era.
Optimism in reinforcement learning and kullback-leibler divergence
S. Filippi, O. Cappé, and A. Garivier · 2010
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Cited alongside, same era.
Algorithms for Reinforcement Learning
Cs. Szepesvári · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Cited alongside, same era.
Later among the works it cites.
Reinforcement learning: An introduction. 2nd edition
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Variance-aware regret bounds for undiscounted reinforcement learning in MDPs
M. S. Talebi and O.-A. Maillard · 2018
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
C. Dann, L. Li, W. Wei, and E. Brunskill · 2019
Later among the works it cites.
Improved analysis of UCRL2B, 2019
R. Fruit, M. Pirotta, and A. Lazaric · 2019
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
C. Jin, Z. Yang, Z. Wang, and M. I. Jordan · 2019
Later among the works it cites.
Bandit algorithms
T. Lattimore and Cs. Szepesvári · 2019
Later among the works it cites.
Online convex optimization in adversarial Markov decision processes
A. Rosenberg and Y. Mansour · 2019
Later among the works it cites.
Worst-case regret bounds for exploration via randomized value functions
D. Russo · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular MDPs
M. Simchowitz and K. G. Jamieson · 2019
Later among the works it cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
L. F. Yang and M. Wang · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
A. Zanette and E. Brunskill · 2019
Later among the works it cites.