Fetching the paper…
Reading the bibliography…
Much of the recent literature on bandit learning focuses on algorithms that aim to converge on an optimal action.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W.R. Thompson · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T.L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Bandit problems with infinitely many arms
D. A. Berry, R. W. Chen, A. Zame, D. C. Heath, and L. A. Shepp · 1997
Earlier work this paper cites.
Multi-armed bandits in metric spaces
R. Kleinberg, A. Slivkins, and E. Upfal · 2008
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Alexandre B Tsybakov · 2008
Earlier work this paper cites.
Algorithms for infinitely many-armed bandits
Y. Wang, J.-Y. Audibert, and R. Munos · 2009
Earlier work this paper cites.
Linearly parameterized bandits
P. Rusmevichientong and J.N. Tsitsiklis · 2010
Earlier work this paper cites.
A modern Bayesian look at the multi-armed bandit
S.L. Scott · 2010
Earlier work this paper cites.
𝒳 \mathcal{X} -armed bandits
S. Bubeck, R. Munos, G. Stoltz, and C. Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
O. Chapelle and L. Li · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Elements of information theory
T.M. Cover and J.A. Thomas · 2012
Cited alongside, same era.
Linear bandits in high dimension and recommendation systems
Y. Deshpande and A. Montanari · 2012
Cited alongside, same era.
The knowledge gradient algorithm for a general class of online learning problems
I.O. Ryzhov, W.B. Powell, and P.I. Frazier · 2012
Cited alongside, same era.
Further optimal regret bounds for Thompson sampling
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
Multi-scale exploration of convex functions and bandit convex optimization
S. Bubeck and R. Eldan · 2015
Later among the works it cites.
On the complexity of best-arm identification in multi-armed bandit models
Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier · 2016
Later among the works it cites.
An information-theoretic analysis of Thompson sampling
D. Russo and B. Van Roy · 2016
Later among the works it cites.
Information directed sampling for stochastic bandits with graph feedback
Fang Liu, Swapna Buccapatnam, and Ness Shroff · 2017
Later among the works it cites.
Thompson sampling for stochastic bandits with graph feedback
Aristide CY Tossou, Christos Dimitrakakis, and Devdatt P Dubhashi · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Two-target algorithms for infinite-armed bandits with Bernoulli rewards
T. Bonald and A. Proutiere · 2013
Cited alongside, same era.
Kullback-Leibler upper confidence bounds for optimal sequential allocation
O. Cappé, A. Garivier, O.-A. Maillard, R. Munos, and G. Stoltz · 2013
Cited alongside, same era.
Learning to optimize via posterior sampling
D. Russo and B. Van Roy · 2014
Cited alongside, same era.
Choosing a good toolkit, I: Formulation, heuristics, and asymptotic properties
A. Francetich and D. M. Kreps
Cited in the paper.
Choosing a good toolkit, II: Simulations and conclusions
A. Francetich and D. M. Kreps
Cited in the paper.
A tutorial on thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al · 2018
Closest in time.
An information-theoretic approach to minimax regret in partial monitoring
Tor Lattimore and Csaba Szepesvári · 2019
Closest in time.
Information-theoretic confidence bounds for reinforcement learning
Xiuyuan Lu and Benjamin Van Roy · 2019
Closest in time.