Fetching the paper…
Reading the bibliography…
The literature on bandit learning and regret analysis has focused on contexts where the goal is to converge on an optimal action in a manner that limits exploration costs.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W.R. Thompson · 1933
Earlier work this paper cites.
Bandit problems with infinitely many arms
D. A. Berry, R. W. Chen, A. Zame, D. C. Heath, and L. A. Shepp · 1997
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Multi-armed bandits in metric spaces
R. Kleinberg, A. Slivkins, and E. Upfal · 2008
Earlier work this paper cites.
Algorithms for infinitely many-armed bandits
Y. Wang, J.-Y. Audibert, and R. Munos · 2009
Earlier work this paper cites.
Linearly parameterized bandits
P. Rusmevichientong and J.N. Tsitsiklis · 2010
Cited alongside, same era.
X-armed bandits
S. Bubeck, R. Munos, G. Stoltz, and C. Szepesvári · 2011
Cited alongside, same era.
Multi-Armed Bandit Allocation Indices
J. Gittins, K. Glazebrook, and R. Weber · 2011
Cited alongside, same era.
Elements of information theory
T.M. Cover and J.A. Thomas · 2012
Cited alongside, same era.
Linear bandits in high dimension and recommendation systems
Y. Deshpande and A. Montanari · 2012
Cited alongside, same era.
Choosing a good toolkit, I: Formulation, heuristics, and asymptotic properties
A. Francetich and D. M. Kreps
Cited in the paper.
Choosing a good toolkit, II: Simulations and conclusions
A. Francetich and D. M. Kreps
Cited in the paper.
The knowledge gradient algorithm for a general class of online learning problems
I.O. Ryzhov, W.B. Powell, and P.I. Frazier · 2012
Later among the works it cites.
Two-target algorithms for infinite-armed bandits with bernoulli rewards
T. Bonald and A. Proutiere · 2013
Later among the works it cites.
Multi-scale exploration of convex functions and bandit convex optimization
S. Bubeck and R. Eldan · 2015
Later among the works it cites.
Bandit convex optimization: T \sqrt{T} regret in one dimension
S. Bubeck, O. Dekel, T. Koren, and Y. Peres · 2015
Later among the works it cites.
An information-theoretic analysis of Thompson sampling
D. Russo and B. Van Roy · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…