Fetching the paper…
Reading the bibliography…
In stochastic bandit problems, a Bayesian policy called Thompson sampling (TS) has recently attracted much attention for its excellent empirical performance.
Thompson][1933]thompson_original Thompson, W. R. (1933) · 1933
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Robbins][1952]robbins Robbins, H. (1952) · 1952
Earlier work this paper cites.
The advanced theory of statistics
Kendall and Stuart][1977]kendall Kendall, M. G., & Stuart, A. (1977) · 1977
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai and Robbins][1985]lai Lai, T. L., & Robbins, H. (1985) · 1985
Earlier work this paper cites.
Optimal adaptive policies for sequential allocation problems
Burnetas and Katehakis][1996]burnetas Burnetas, A. N., & Katehakis, M. N. (1996) · 1996
Earlier work this paper cites.
Large deviations techniques and applications
Dembo and Zeitouni][1998]LDP Dembo, A., & Zeitouni, O. (1998) · 1998
Cited alongside, same era.
The Bayesian choice
Robert][2001]bayes_robert Robert, C. P. (2001) · 2001
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Auer et al.][2002]ucb Auer, P., Cesa-Bianchi, N., & Fischer, P. (2002) · 2002
Cited alongside, same era.
An asymptotically optimal bandit algorithm for bounded support models
Honda and Takemura][2010]honda_colt Honda, J., & Takemura, A. (2010) · 2010
Cited alongside, same era.
NIST handbook of mathematical functions
Olver et al.][2010]nist_handbook Olver, F. W., Lozier, D. W., Boisvert, R. F., & Clark, C. W. (2010) · 2010
Cited alongside, same era.
On bayesian upper confidence bounds for bandit problems
Kaufmann et al.][2012a]bayes_ucb Kaufmann, E., Cappé, O., & Garivier, A. (2012a)
Cited in the paper.
An empirical evaluation of Thompson sampling
Chapelle and Li][2012]thompson_empirical Chapelle, O., & Li, L. (2012) · 2011
Later among the works it cites.
The KL-UCB algorithm for bounded stochastic bandits and beyond
Garivier and Cappé][2011]kl_ucb Garivier, A., & Cappé, O. (2011) · 2011
Later among the works it cites.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal and Goyal][2012]thompson_log Agrawal, S., & Goyal, N. (2012) · 2012
Later among the works it cites.
Thompson sampling for 1-dimensional exponential family bandits
Korda et al.][2013]thompson_exponential Korda, N., Kaufmann, E., & Munos, R. (2013) · 2013
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thompson sampling: an asymptotically optimal finite-time analysis
Kaufmann et al.][2012b]thompson Kaufmann, E., Korda, N., & Munos, R. (2012b)
Cited in the paper.