Fetching the paper…
Reading the bibliography…
We consider the problem of learning to optimize an unknown Markov decision process (MDP).
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William Thompson · 1933
Earlier work this paper cites.
A theory of the learnable
Leslie G Valiant · 1984
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm
Nick Littlestone · 1988
Earlier work this paper cites.
Dynamic programming and optimal control
Dimitri Bertsekas · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Earlier work this paper cites.
Optimal adaptive policies for Markov decision processes
Apostolos Burnetas and Michael Katehakis · 1997
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Malcom Strens · 2000
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Cited alongside, same era.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen Brafman and Moshe Tennenholtz · 2003
Cited alongside, same era.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2003
Cited alongside, same era.
Pac model-free reinforcement learning
Alexander Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael Littman · 2006
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Knows what it knows: a framework for self-aware learning
Lihong Li, Michael L Littman, Thomas J Walsh, and Alexander L Strehl · 2011
Cited alongside, same era.
Efficient reinforcement learning for high dimensional linear quadratic systems
Morteza Ibrahimi, Adel Javanmard, and Benjamin Van Roy · 2012
Later among the works it cites.
Online regret bounds for undiscounted continuous reinforcement learning
Ronald Ortner, Daniil Ryabko, et al · 2012
Later among the works it cites.
(More) Efficient Reinforcement Learning via Posterior Sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Later among the works it cites.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2013
Later among the works it cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Later among the works it cites.
Near-optimal regret bounds for reinforcement learning in factored MDPs
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X-armed bandits
Sébastien Bubeck, Rémi Munos, Gilles Stoltz, and Csaba Szepesvári · 2011
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Yassin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Cited alongside, same era.
Ian Osband and Benjamin Van Roy · 2014
Closest in time.
Generalization and exploration via randomized value functions
Benjamin Van Roy and Zheng Wen · 2014
Closest in time.