Fetching the paper…
Reading the bibliography…
Statistical performance bounds for reinforcement learning (RL) algorithms can be critical for high-stakes applications like healthcare.
Using upper confidence bounds for online learning
Peter Auer · 2000
Earlier work this paper cites.
Inequalities for the L 1 Deviation of the Empirical Distribution
Tsachy Weissman, Erik Ordentlich, Gadiel Seroussi, Sergio Verdu, and Marcelo J Weinberger · 2003
Earlier work this paper cites.
Online Regret Bounds for a New Reinforcement Learning Algorithm
Peter Auer and Ronald Ortner · 2005
Earlier work this paper cites.
Concentration inequalities and model selection
Pascal Massart · 2007
Earlier work this paper cites.
An analysis of model-based Interval Estimation for Markov Decision Processes
Alexander L. Strehl and Michael L. Littman · 2008
Earlier work this paper cites.
Reinforcement Learning in Finite MDPs : PAC Analysis
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Earlier work this paper cites.
Exploration-exploitation tradeoff using variance estimates in multi-armed bandits
Jean Yves Audibert, Rémi Munos, and Csaba Szepesvári · 2009
Earlier work this paper cites.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Peter L. Bartlett and a. Tewari · 2009
Earlier work this paper cites.
Near-optimal Regret Bounds for Reinorcement Learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Istvàn Szita and Csaba Szepesvári · 2010
Cited alongside, same era.
Probability - Theory and Examples
Rick Durrett · 2010
Cited alongside, same era.
Knows what it knows: A framework for self-aware learning
Lihong Li, Michael L. Littman, Thomas J. Walsh, and Alexander L. Strehl · 2011
Cited alongside, same era.
The KL-UCB Algorithm for Bounded Stochastic Bandits and Beyond
Aurelien Garivier and Olivier Cappe · 2011
Cited alongside, same era.
Information-theoretic regret bounds for Gaussian process optimization in the bandit setting
Niranjan Srinivas, Andreas Krause, Sham M. Kakade, and Matthias W. Seeger · 2012
Cited alongside, same era.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Taming the Monster: A Fast and Simple Algorithm for Contextual Bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert E. Schapire · 2014
Later among the works it cites.
How to discount deep reinforcement learning: Towards new dynamic strategies
Vincent François-Lavet, Raphaël Fonteneau, and Damien Ernst · 2015
Later among the works it cites.
Sample Complexity of Episodic Fixed-Horizon Reinforcement Learning
Christoph Dann and Emma Brunskill · 2015
Later among the works it cites.
Efficient PAC-optimal Exploration in Concurrent , Continuous State MDPs with Delayed Updates
Jason Pazis and Ronald Parr · 2016
Later among the works it cites.
Sequential Nonparametric Testing with the Law of the Iterated Logarithm
Akshay Balsubramani and Aaditya Ramdas · 2016
Later among the works it cites.
On Explore-Then-Commit Strategies
Aurélien Garivier, Emilie Kaufmann, and Tor Lattimore · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
lil’ UCB : An Optimal Exploration Algorithm for Multi-Armed Bandits
Kevin Jamieson, Matthew Malloy, Robert Nowak, and Sébastien Bubeck · 2013
Cited alongside, same era.
Concentration Inequalities - A Nonasymptotic Theory of Independence
Stephane Boucheron, Gabor Lugosi, and Pascal Massart · 2013
Cited alongside, same era.
Near-optimal PAC bounds for discounted MDPs
Tor Lattimore and Marcus Hutter · 2014
Cited alongside, same era.
Later among the works it cites.
Contextual Decision Processes with Low Bellman Rank are PAC-Learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Closest in time.
Minimax Regret Bounds for Reinforcement Learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Closest in time.