Fetching the paper…
Reading the bibliography…
It is well known that in stochastic multi-armed bandits (MAB), the sample mean of an arm is typically not an unbiased estimator of its true mean.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Remarks on a stopping time
Norman Starr and Michael B Woodroofe · 1968
Earlier work this paper cites.
Statistical methods related to the law of the iterated logarithm
Herbert Robbins · 1970
Earlier work this paper cites.
Estimation following sequential tests
David Siegmund · 1978
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
Jean-Yves Audibert and Sébastien Bubeck · 2009
Earlier work this paper cites.
Stopped random walks
Allan Gut · 2009
Cited alongside, same era.
The KL-UCB algorithm for bounded stochastic bandits and beyond
Aurélien Garivier and Olivier Cappé · 2011
Cited alongside, same era.
Analysis of thompson sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal · 2012
Cited alongside, same era.
Pac subset selection in stochastic multi-armed bandits
Shivaram Kalyanakrishnan, Ambuj Tewari, Peter Auer, and Peter Stone · 2012
Cited alongside, same era.
Thompson sampling: An asymptotically optimal finite-time analysis
Emilie Kaufmann, Nathaniel Korda, and Rémi Munos · 2012
Cited alongside, same era.
Estimation bias in multi-armed bandit algorithms for search advertising
Min Xu, Tao Qin, and Tie-Yan Liu · 2013
Cited alongside, same era.
Optimal best arm identification with fixed confidence
Aurélien Garivier and Emilie Kaufmann · 2016
Later among the works it cites.
Simple bayesian algorithms for best arm identification
Daniel Russo · 2016
Later among the works it cites.
Unbiased estimation for response adaptive clinical trials
Jack Bowden and Lorenzo Trippa · 2017
Later among the works it cites.
Uniform, nonparametric, non-asymptotic confidence sequences
Steven R Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon · 2018
Later among the works it cites.
Why adaptively collected data have negative bias and how to correct for it
Xinkun Nie, Xiaoying Tian, Jonathan Taylor, and James Zou · 2018
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
lil’ UCB: An Optimal Exploration Algorithm for Multi-Armed Bandits
Kevin Jamieson, Matthew Malloy, Robert Nowak, and Sébastien Bubeck · 2014
Cited alongside, same era.
Multi-armed bandit models for the optimal design of clinical trials: benefits and challenges
Sofía S Villar, Jack Bowden, and James Wason · 2015
Cited alongside, same era.
Closest in time.
On the bias, risk and consistency of sample means in multi-armed bandits
Jaehyeok Shin, Aaditya Ramdas, and Alessandro Rinaldo · 2019
Closest in time.