Fetching the paper…
Reading the bibliography…
I present the first algorithm for stochastic finite-armed bandits that simultaneously enjoys order-optimal problem-dependent regret and worst-case regret.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William Thompson · 1933
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
On sequential designs for maximizing the sum of n observations
Russell N Bradt, SM Johnson, and Samuel Karlin · 1956
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
John Gittins · 1979
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Sample mean based index policies with O(log n) regret for the multi-armed bandit problem
Rajeev Agrawal · 1995
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 1995
Earlier work this paper cites.
Sequential choice from several populations
Michael N Katehakis and Herbert Robbins · 1995
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicoló Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Tuning bandit algorithms in stochastic environments
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvári · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Cited alongside, same era.
Introduction to nonparametric estimation
Alexandre B Tsybakov · 2008
Cited alongside, same era.
Minimax policies for adversarial and stochastic bandits
Jean-Yves Audibert and Sébastien Bubeck · 2009
Cited alongside, same era.
Pure exploration in multi-armed bandits problems
Sébastien Bubeck, Rémi Munos, and Gilles Stoltz · 2009
Cited alongside, same era.
Best arm identification in multi-armed bandits
Jean-Yves Audibert and Sébastien Bubeck · 2010
Cited alongside, same era.
UCB revisited: Improved regret bounds for the stochastic multi-armed bandit problem
Peter Auer and Ronald Ortner · 2010
Cited alongside, same era.
The KL-UCB algorithm for bounded stochastic bandits and beyond
Aurélien Garivier · 2011
Later among the works it cites.
Finite-time analysis of multi-armed bandits problems with Kullback-Leibler divergences
Odalric-Ambrym Maillard, Rémi Munos, and Gilles Stoltz · 2011
Later among the works it cites.
Computing a classic index for finite-horizon bandits
José Niño-Mora · 2011
Later among the works it cites.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Later among the works it cites.
Bandit theory meets compressed sensing for high dimensional stochastic linear bandit
Alexandra Carpentier and Rémi Munos · 2012
Later among the works it cites.
Bandits with heavy tail
Sebastian Bubeck, Nicolo Cesa-Bianchi, and Gábor Lugosi · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bandits games and clustering foundations
Sébastien Bubeck · 2010
Cited alongside, same era.
Linearly parameterized bandits
Paat Rusmevichientong and John N Tsitsiklis · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Csaba Szepesvári, and David Tax · 2011
Cited alongside, same era.
An empirical evaluation of thompson sampling
Olivier Chapelle and Lihong Li · 2011
Cited alongside, same era.
Further optimal regret bounds for thompson sampling
Shipra Agrawal and Navin Goyal
Cited in the paper.
Analysis of thompson sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal
Cited in the paper.
Kullback–Leibler upper confidence bounds for optimal sequential allocation
Olivier Cappé, Aurélien Garivier, Odalric-Ambrym Maillard, Rémi Munos, and Gilles Stoltz · 2013
Later among the works it cites.
Thompson sampling for 1-dimensional exponential family bandits
Nathaniel Korda, Emilie Kaufmann, and Rémi Munos · 2013
Later among the works it cites.
lil’UCB: An optimal exploration algorithm for multi-armed bandits
Kevin Jamieson, Matthew Malloy, Robert Nowak, and Sébastien Bubeck · 2014
Later among the works it cites.
Regret analysis of the finite-horizon Gittins index strategy for multi-armed bandits
Tor Lattimore · 2015
Closest in time.