Fetching the paper…
Reading the bibliography…
We provide an information-theoretic analysis of Thompson sampling that applies across a broad range of online optimization problems in which a decision-maker must learn from partial feedback.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W.R. Thompson · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T.L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Convex optimization
S.P. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
The price of bandit information for online optimization
V. Dani, S.M. Kakade, and T.P. Hayes · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
V. Dani, T.P. Hayes, and S.M. Kakade · 2008
Earlier work this paper cites.
Web-scale Bayesian click-through rate prediction for sponsored search advertising in Microsoft’s Bing search engine
T. Graepel, J.Q. Candela, T. Borchert, and R. Herbrich · 2010
Earlier work this paper cites.
Linearly parameterized bandits
P. Rusmevichientong and J.N. Tsitsiklis · 2010
Earlier work this paper cites.
A modern Bayesian look at the multi-armed bandit
S.L. Scott · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
O. Chapelle and L. Li · 2011
Cited alongside, same era.
Multi-Armed Bandit Allocation Indices
J. Gittins, K. Glazebrook, and R. Weber · 2011
Cited alongside, same era.
Entropy and information theory
R.M. Gray · 2011
Cited alongside, same era.
Information-theoretic regret bounds for Gaussian process optimization in the bandit setting
N. Srinivas, A. Krause, S.M. Kakade, and M. Seeger · 2011
Cited alongside, same era.
Analysis of Thompson sampling for the multi-armed bandit problem
S. Agrawal and N. Goyal · 2012
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
The knowledge gradient algorithm for a general class of online learning problems
I.O. Ryzhov, W.B. Powell, and P.I. Frazier · 2012
Later among the works it cites.
Regret in online combinatorial optimization
J.-Y. Audibert, S. Bubeck, and G. Lugosi · 2013
Later among the works it cites.
Prior-free and prior-dependent regret bounds for Thompson sampling
S. Bubeck and C.-Y. Liu · 2013
Later among the works it cites.
Thompson sampling for complex bandit problems
A. Gopalan, S. Mannor, and Y. Mansour · 2013
Later among the works it cites.
Thompson sampling for one-dimensional exponential family bandits
N. Korda, E. Kaufmann, and R. Munos · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Elements of information theory
T.M. Cover and J.A. Thomas · 2012
Cited alongside, same era.
Thompson sampling: an asymptotically optimal finite time analysis
E. Kauffmann, N. Korda, and R. Munos · 2012
Cited alongside, same era.
Optimistic Bayesian sampling in contextual-bandit problems
B.C. May, N. Korda, A. Lee, and D.S. Leslie · 2012
Cited alongside, same era.
Further optimal regret bounds for thompson sampling
S. Agrawal and N. Goyal
Cited in the paper.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal
Cited in the paper.
L. Li · 2013
Later among the works it cites.
Learning to optimize via posterior sampling
D. Russo and B. Van Roy · 2013
Later among the works it cites.
Automatic ad format selection via contextual bandits
L. Tang, R. Rosales, A. Singh, and D. Agarwal · 2013
Later among the works it cites.
Overview of content experiments: Multi-armed bandit experiments, 2014
S.L. Scott · 2014
Closest in time.