Fetching the paper…
Reading the bibliography…
Recently, there has been significant progress in understanding reinforcement learning in discounted infinite-horizon Markov decision processes (MDPs) by deriving tight sample complexity bounds.
The Variance of Markov Decision Processes
Matthew J Sobel · 1982
Earlier work this paper cites.
Efficient reinforcement learning
Claude-Nicolas Fiechter · 1994
Earlier work this paper cites.
Expected Mistake Bound Model for On-Line Reinforcement Learning
Claude-Nicolas Fiechter · 1997
Earlier work this paper cites.
Finite-Sample Convergence Rates for Q-Learning and Indirect Algorithms
Michael J Kearns and Satinder P Singh · 1999
Earlier work this paper cites.
R-MAX – A General Polynomail Time Algorithm for Near-Optimal Reinforcement Learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Sham M. Kakade · 2003
Earlier work this paper cites.
The Sample Complexity of Exploration in the Multi-Armed Bandit Problem
Shie Mannor and John N Tsitsiklis · 2004
Cited alongside, same era.
Online Regret Bounds for a New Reinforcement Learning Algorithm
Peter Auer and Ronald Ortner · 2005
Cited alongside, same era.
Concentration Inequalities and Martingale Inequalities: A Survey
Fan Chung and Linyuan Lu · 2006
Cited alongside, same era.
Efficient PAC learning for episodic tasks with acyclic state spaces
Spyros Reveliotis and Theologos Bountourelis · 2007
Cited alongside, same era.
An analysis of model-based Interval Estimation for Markov Decision Processes
Alexander L. Strehl and Michael L. Littman · 2008
Cited alongside, same era.
Near-Bayesian exploration in polynomial time
J Zico Kolter and Andrew Y Ng · 2009
Reinforcement Learning in Finite MDPs : PAC Analysis
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Later among the works it cites.
Empirical Bernstein Bounds and Sample-Variance Penalization
Andreas Maurer and Massimiliano Pontil · 2009
Later among the works it cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Istvàn Szita and Csaba Szepesvári · 2010
Later among the works it cites.
PAC bounds for discounted MDPs
Tor Lattimore and Marcus Hutter · 2012
Later among the works it cites.
On the Sample Complexity of Reinforcement Learning with a Generative Model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J. Kappen · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
PAC Model-Free Reinforcement Learning
Alexander L. Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L. Littman
Cited in the paper.
Incremental Model-based Learners With Formal Learning-Time Guarantees
Alexander L Strehl, Lihong Li, and Michael L Littman
Cited in the paper.
Near-optimal Regret Bounds for Reinforcement Learning
Thomas Jaksch, Ronald Ortner, and Peter Auer
Cited in the paper.
Near-optimal Regret Bounds for Reinforcement Learning
Thomas Jaksch, Ronald Ortner, and Peter Auer
Cited in the paper.