Fetching the paper…
Reading the bibliography…
The stochastic multi-armed bandit model is a simple abstraction that has proven useful in many different contexts in statistics and machine learning.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W.R. Thompson · 1933
Earlier work this paper cites.
Sequential tests of statistical hypotheses
A. Wald · 1945
Earlier work this paper cites.
Some aspects of the sequential design of experiments
H. Robbins · 1952
Earlier work this paper cites.
A sequential procedure for selecting the population with the largest mean from k normal populations
E. Paulson · 1964
Earlier work this paper cites.
Sequential identification and ranking procedures
Robert Bechhofer, Jack Kiefer, and Milton Sobel · 1968
Earlier work this paper cites.
Statistical methods related to the law of the iterated logarithm
H. Robbins · 1970
Earlier work this paper cites.
Asymptotically optimal procedures for sequential adaptive selection of the best of several normal means
C. Jennison, I.M. Johnstone, and B.W. Turnbull · 1982
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T.L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Sequential analysis
D. Siegmund · 1985
Earlier work this paper cites.
Optimal adaptive policies for sequential allocation problems
A.N Burnetas and M. Katehakis · 1996
Cited alongside, same era.
The Racing algorithm: Model selection for lazy learners
O. Maron and A. Moore · 1997
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Cited alongside, same era.
The sample complexity of exploration in the multi-armed bandit problem
S. Mannor and J. Tsitsiklis · 2004
Cited alongside, same era.
Elements of information theory (2nd Edition)
T. Cover and J. Thomas · 2006
Cited alongside, same era.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
E. Even-Dar, S. Mannor, and Y. Mansour · 2006
Cited alongside, same era.
An asymptotically optimal policy for finite support models in the multiarmed bandit problem
J. Honda and A. Takemura · 2011
Later among the works it cites.
Best arm identification: a unified approach to fixed budget and fixed confidence
V. Gabillon, M. Ghavamzadeh, and A. Lazaric · 2012
Later among the works it cites.
PAC subset selection in stochastic multi-armed bandits
S. Kalyanakrishnan, A. Tewari, P. Auer, and P. Stone · 2012
Later among the works it cites.
Further optimal regret bounds for Thompson Sampling
S. Agrawal and N. Goyal · 2013
Later among the works it cites.
Kullback-Leibler upper confidence bounds for optimal sequential allocation
O. Cappé, A. Garivier, O-A. Maillard, R. Munos, and G. Stoltz · 2013
Later among the works it cites.
Almost optimal exploration in multi-armed bandits
Z. Karnin, T. Koren, and O. Somekh · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hoeffding and Bernstein races for selecting policies in evolutionary direct policy search
V. Heidrich-Meisner and C. Igel · 2009
Cited alongside, same era.
Best arm identification in multi-armed bandits
J-Y. Audibert, S. Bubeck, and R. Munos · 2010
Cited alongside, same era.
Pure exploration in finitely armed and continuous armed bandits
S. Bubeck, R. Munos, and G. Stoltz · 2011
Cited alongside, same era.
Bounded regret in stochastic multi-armed bandits
S. Bubeck, V. Perchet, and P. Rigollet
Cited in the paper.
Multiple identifications in multi-armed bandits
S. Bubeck, T. Wang, and N. Viswanathan
Cited in the paper.
On Bayesian upper-confidence bounds for bandit problems
E. Kaufmann, A. Garivier, and O. Cappé
Cited in the paper.
Later among the works it cites.
Information complexity in bandit subset selection
E. Kaufmann and S. Kalyanakrishnan · 2013
Later among the works it cites.
lil’UCB: an optimal exploration algorithm for multi-armed bandits
K. Jamieson, M. Malloy, R. Nowak, and S. Bubeck · 2014
Closest in time.