Fetching the paper…
Reading the bibliography…
Thompson Sampling is one of the oldest heuristics for multi-armed bandit problems.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Multi-armed Bandit Allocation Indices
J. C. Gittins · 1989
Earlier work this paper cites.
Exploration and Inference in Learning from Reinforcement
J. Wyatt · 1997
Earlier work this paper cites.
A Bayesian Framework for Reinforcement Learning
M. J. A. Strens · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Dual weak pigeonhole principle, Boolean complexity, and derandomization
E. Jeřábek · 2004
Earlier work this paper cites.
Minimax Policies for Adversarial and Stochastic Bandits
J.-Y. Audibert and S. Bubeck · 2009
Cited alongside, same era.
Web-Scale Bayesian Click-Through rate Prediction for Sponsored Search Advertising in Microsoft’s Bing Search Engine
T. Graepel, J. Q. Candela, T. Borchert, and R. Herbrich · 2010
Cited alongside, same era.
Solving Two-Armed Bernoulli Bandit Problems Using a Bayesian Learning Automaton
O.-C. Granmo · 2010
Cited alongside, same era.
Linearly parametrized bandits
P. A. Ortega and D. A. Braun · 2010
Cited alongside, same era.
A modern Bayesian look at the multi-armed bandit
S. Scott · 2010
Cited alongside, same era.
An Empirical Evaluation of Thompson Sampling
O. Chapelle and L. Li · 2011
Cited alongside, same era.
Finite-time analysis of multi-armed bandits problems with Kullback-Leibler divergences
O.-A. Maillard, R. Munos, and G. Stoltz · 2011
Later among the works it cites.
Analysis of Thompson Sampling for the Multi-armed Bandit Problem
S. Agrawal and N. Goyal · 2012
Closest in time.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawan and N. Goyal · 2012
Closest in time.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
S. Bubeck and N. Cesa-Bianchi · 2012
Closest in time.
Open Problem: Regret Bounds for Thompson Sampling
O. Chapelle and L. Li · 2012
Closest in time.
On Bayesian Upper Confidence Bounds for Bandit Problems
E. Kaufmann, O. Cappé, and A. Garivier · 2012
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The KL-UCB Algorithm for Bounded Stochastic Bandits and Beyond
A. Garivier and O. Cappé · 2011
Cited alongside, same era.
Optimistic Bayesian sampling in contextual-bandit problems
B. C. May, N. Korda, A. Lee, and D. S. Leslie
Cited in the paper.
Simulation studies in optimistic Bayesian sampling in contextual-bandit problems
B. C. May and D. S. Leslie
Cited in the paper.
Thompson Sampling: An Optimal Finite Time Analysis
E. Kaufmann, N. Korda, and R. Munos · 2012
Closest in time.