Fetching the paper…
Reading the bibliography…
A/B testing refers to the task of determining the best option among two alternatives that yield random outcomes.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W.R. Thompson · 1933
Earlier work this paper cites.
Sequential tests of statistical hypotheses
A. Wald · 1945
Earlier work this paper cites.
Some aspects of the sequential design of experiments
H. Robbins · 1952
Earlier work this paper cites.
Sequential identification and ranking procedures
Robert Bechhofer, Jack Kiefer, and Milton Sobel · 1968
Earlier work this paper cites.
Statistical Methods Related to the law of the iterated logarithm
H. Robbins · 1970
Earlier work this paper cites.
Asymptotically optimal procedures for sequential adaptive selection of the best of several normal means
C. Jennison, I.M. Johnstone, and B.W. Turnbull · 1982
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T.L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Sequential Analysis
D. Siegmund · 1985
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Cited alongside, same era.
The Sample Complexity of Exploration in the Multi-Armed Bandit Problem
S. Mannor and J. Tsitsiklis · 2004
Cited alongside, same era.
Elements of Information Theory (2nd Edition)
T. Cover and J. Thomas · 2006
Cited alongside, same era.
Action Elimination and Stopping Conditions for the Multi-Armed Bandit and Reinforcement Learning Problems
E. Even-Dar, S. Mannor, and Y. Mansour · 2006
Cited alongside, same era.
Best Arm Identification in Multi-armed Bandits
J-Y. Audibert, S. Bubeck, and R. Munos · 2010
Cited alongside, same era.
Pure Exploration in Finitely Armed and Continuous Armed Bandits
S. Bubeck, R. Munos, and G. Stoltz · 2011
Cited alongside, same era.
Best Arm Identification: A Unified Approach to Fixed Budget and Fixed Confidence
V. Gabillon, M. Ghavamzadeh, and A. Lazaric · 2012
Later among the works it cites.
PAC subset selection in stochastic multi-armed bandits
S. Kalyanakrishnan, A. Tewari, P. Auer, and P. Stone · 2012
Later among the works it cites.
Further Optimal Regret Bounds for Thompson Sampling
S. Agrawal and N. Goyal · 2013
Later among the works it cites.
Multiple Identifications in multi-armed bandits
S. Bubeck, T. Wang, and N. Viswanathan · 2013
Later among the works it cites.
Kullback-Leibler upper confidence bounds for optimal sequential allocation
O. Cappé, A. Garivier, O-A. Maillard, R. Munos, and G. Stoltz · 2013
Later among the works it cites.
lil’UCB: an optimal exploration algorithm for multi-armed bandits
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An asymptotically optimal policy for finite support models in the multiarmed bandit problem
J. Honda and A. Takemura · 2011
Cited alongside, same era.
On Bayesian Upper-Confidence Bounds for Bandit Problems
E. Kaufmann, A. Garivier, and O. Cappé
Cited in the paper.
Thompson Sampling : an Asymptotically Optimal Finite-Time Analysis
E. Kaufmann, N. Korda, and R. Munos
Cited in the paper.
K. Jamieson, M. Malloy, R. Nowak, and S. Bubeck · 2013
Later among the works it cites.
Information complexity in bandit subset selection
E. Kaufmann and S. Kalyanakrishnan · 2013
Later among the works it cites.