Fetching the paper…
Reading the bibliography…
The empirically successful Thompson Sampling algorithm for stochastic bandits has drawn much interest in understanding its theoretical properties.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. Thompson · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and H. Robbins · 1985
Earlier work this paper cites.
The non-stochastic multi-armed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. Schapire · 2002
Earlier work this paper cites.
Prediction, Learning, and Games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Web-scale Bayesian click-through rate prediction for sponsored search advertising in Microsoft’s Bing search engine
T. Graepel, J. Quinonero Candela, T. Borchert, and R. Herbrich · 2010
Earlier work this paper cites.
A modern Bayesian look at the multi-armed bandit
S. L. Scott · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and Cs. Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
O. Chapelle and L. Li · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
W. Chu, L. Li, L. Reyzin, and R.E. Schapire · 2011
Earlier work this paper cites.
Analysis of Thompson sampling for the multi-armed bandit problem
S. Agrawal and N. Goyal · 2012
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Thompson sampling: An asymptotically optimal finite-time analysis
E. Kaufmann, N. Korda, and R. Munos · 2012
Cited alongside, same era.
Optimistic Bayesian sampling in contextual-bandit problems
B. C. May, N. Korda, A. Lee, and D. S. Leslie · 2012
Cited alongside, same era.
Further optimal regret bounds for Thompson sampling
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
Sequential Experimentation in Clinical Trials: Design and Analysis
J. Bartroff, T. L. Lai, and M.-C. Shih · 2013
Cited alongside, same era.
Thompson sampling for complex online problems
A. Gopalan, S. Mannor, and Y. Mansour · 2014
Later among the works it cites.
Stochastic regret minimization via Thompson sampling
S. Guha and K. Munagala · 2014
Later among the works it cites.
Optimality of Thompson sampling for gaussian bandits depends on priors
J. Honda and A. Takemura · 2014
Later among the works it cites.
Learning to optimize via posterior sampling
D. Russo and B. Van Roy · 2014
Later among the works it cites.
Optimal regret analysis of Thompson sampling in stochastic multi-armed bandit problem with multiple plays
J. Komiyama, J. Honda, and H. Nakagawa · 2015
Closest in time.
The Pareto regret frontier for bandits
T. Lattimore · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Prior-free and prior-dependent regret bounds for Thompson sampling
S. Bubeck and C.Y. Liu · 2013
Cited alongside, same era.
Approximation algorithms for Bayesian multi-armed bandit problems
S. Guha and K. Munagala · 2013
Cited alongside, same era.
Generalized Thompson sampling for contextual bandits
L. Li · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. E. Schapire · 2014
Cited alongside, same era.
C.Y. Liu and L. Li · 2015
Closest in time.
Thompson sampling for budgeted multi-armed bandits
Y. Xia, H. Li, T. Qin, and N. Yu ans T.-Y. Liu · 2015
Closest in time.
Towards optimal algorithms for prediction with expert advice
N. Gravin, Y. Peres, and B. Sivan · 2016
Closest in time.
An information-theoretic analysis of Thompson sampling
D. Russo and B. Van Roy · 2016
Closest in time.