Fetching the paper…
Reading the bibliography…
An online reinforcement learning algorithm is anytime if it does not need to know in advance the horizon T of the experiment.
On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples
W. R. Thompson · 1933
Earlier work this paper cites.
Some Aspects of the Sequential Design of Experiments
H. Robbins · 1952
Earlier work this paper cites.
Asymptotically Efficient Adaptive Allocation Rules
T. L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Gambling in a Rigged Casino: The Adversarial Multi-Armed Bandit Problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. Schapire · 1995
Earlier work this paper cites.
Prediction, Learning, and Games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
J-Y. Audibert and S. Bubeck · 2009
Earlier work this paper cites.
Multi-Armed Bandit Based Policies for Cognitive Radio’s Decision Making Issues
W. Jouini, D. Ernst, C. Moy, and J. Palicot · 2009
Earlier work this paper cites.
UCB Revisited: Improved Regret Bounds For The Stochastic Multi-Armed Bandit Problem
P. Auer and R. Ortner · 2010
Earlier work this paper cites.
An Asymptotically Optimal Bandit Algorithm for Bounded Support Models
J. Honda and A. Takemura · 2010
Cited alongside, same era.
A Contextual-Bandit Approach to Personalized News Article Recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Cited alongside, same era.
An Empirical Evaluation of Thompson Sampling
O. Chapelle and L. Li · 2011
Cited alongside, same era.
Analysis of Thompson sampling for the Multi-Armed Bandit problem
S. Agrawal and N. Goyal · 2012
Cited alongside, same era.
Regret Analysis of Stochastic and Non-Stochastic Multi-Armed Bandit Problems
S. Bubeck, N. Cesa-Bianchi, et al · 2012
Cited alongside, same era.
Thompson Sampling: an Asymptotically Optimal Finite-Time Analysis , pages 199–213
E. Kaufmann, N. Korda, and R. Munos · 2012
Cited alongside, same era.
On the Complexity of A/B Testing
E. Kaufmann, O. Cappé, and A. Garivier · 2014
Later among the works it cites.
Anytime Optimal Algorithms In Stochastic Multi Armed Bandits
R. Degenne and V. Perchet · 2016
Later among the works it cites.
On Explore-Then-Commit Strategies
A. Garivier, E. Kaufmann, and T. Lattimore · 2016
Later among the works it cites.
Regret Analysis Of The Finite Horizon Gittins Index Strategy For Multi Armed Bandits
T. Lattimore · 2016
Later among the works it cites.
A Minimax and Asymptotically Optimal Algorithm for Stochastic Bandits
P. Ménard and A. Garivier · 2017
Later among the works it cites.
An Improved Parametrization and Analysis of the EXP3++ Algorithm for Stochastic and Adversarial Bandits
Y. Seldin and G. Lugosi · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Risk-Aversion In Multi-Armed Bandits
A. Sani, A. Lazaric, and R. Munos · 2012
Cited alongside, same era.
Kullback-Leibler upper confidence bounds for optimal sequential allocation
O. Cappé, A. Garivier, O-A. Maillard, R. Munos, and G. Stoltz · 2013
Cited alongside, same era.
Finite-time Analysis of the Multi-armed Bandit Problem
P. Auer, N. Cesa-Bianchi, and P. Fischer
Cited in the paper.
The Nonstochastic Multiarmed Bandit Problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. Schapire
Cited in the paper.
A framework for Multi-A(rmed)/B(andit) Testing with Online FDR Control
F. Yang, A. Ramdas, K. Jamieson, and M. Wainwright · 2017
Later among the works it cites.
Stochastic Multi-Armed Bandits in Constant Space
D. Liau, E. Price, Z. Song, and G. Yang · 2018
Closest in time.