Fetching the paper…
Reading the bibliography…
The problem of reinforcement learning in an unknown and discrete Markov Decision Process (MDP) under the average-reward criterion is considered, when the learner interacts with the system in a single stream of observations, starting from an initial state without any reset.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Optimal adaptive policies for M
Apostolos N. Burnetas and Michael N. Katehakis · 1997
Earlier work this paper cites.
Asymptotically efficient adaptive choice of control laws in controlled m
Todd L. Graves and Tze Leung Lai · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Robust control of M
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for M
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible MDP
Ambuj Tewari and Peter L. Bartlett · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2009
Cited alongside, same era.
Stratégies optimistes en apprentissage par renforcement
Sarah Filippi · 2010
Cited alongside, same era.
Optimism in reinforcement learning and K
Sarah Filippi, Olivier Cappé, and Aurélien Garivier · 2010
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Some bounds for the logarithmic function
Flemming Topsøe · 2010
Cited alongside, same era.
Concentration inequalities: A nonasymptotic theory of independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Cited alongside, same era.
How hard is my MDP
Odalric-Ambrym Maillard, Timothy A. Mann, and Shie Mannor · 2014
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L. Puterman · 2014
Later among the works it cites.
Explore first, exploit next: The true shape of regret in bandit problems
Aurélien Garivier, Pierre Ménard, and Gilles Stoltz · 2016
Later among the works it cites.
Optimistic posterior sampling for reinforcement learning: Worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Later among the works it cites.
Unifying PAC
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Dan Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Learning unknown M
Yi Ouyang, Mukul Gagrani, Ashutosh Nayyar, and Rahul Jain · 2017
Later among the works it cites.