Fetching the paper…
Reading the bibliography…
This technical note presents a new approach to carrying out the kind of exploration achieved by Thompson sampling, but without explicitly maintaining or sampling from posterior distributions.
Bootstrap methods: another look at the jackknife
Bradley Efron · 1979
Earlier work this paper cites.
The bayesian bootstrap
Donald B Rubin et al · 1981
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Cited alongside, same era.
Further optimal regret bounds for Thompson sampling
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
(More) Efficient Reinforcement Learning via Posterior Sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Thompson sampling with the online bootstrap
Dean Eckles and Maurits Kaptein · 2014
Cited alongside, same era.
Sub-sampling for multi-armed bandits
Akram Baransi, Odalric-Ambrym Maillard, and Shie Mannor · 2014
Later among the works it cites.
Near-optimal regret bounds for reinforcement learning in factored MDPs
Ian Osband and Benjamin Van Roy · 2014
Later among the works it cites.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Later among the works it cites.
Generalization and exploration via randomized value functions
Benjamin Van Roy and Zheng Wen · 2014
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…