Fetching the paper…
Reading the bibliography…
We define a novel family of algorithms for the adversarial multi-armed bandit problem, and provide a simple analysis technique based on convex smoothing.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Approximation to bayes risk in repeated play
J. Hannan · 1957
Earlier work this paper cites.
Stochastic optimization problems with nondifferentiable cost functionals
Dimitri P. Bertsekas · 1973
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Possible generalization of boltzmann-gibbs statistics
Constantino Tsallis · 1988
Earlier work this paper cites.
The weighted majority algorithm
Nick Littlestone and Manfred K. Warmuth · 1994
Earlier work this paper cites.
Sub-hessians, super-hessians and conjugation
Jean-Paul Penot · 1994
Earlier work this paper cites.
Quantitative methods in the planning of pharmaceutical research
John Gittins · 1996
Earlier work this paper cites.
Modelling Extremal Events: For Insurance and Finance
P. Embrechts, C. Klüppelberg, and T. Mikosch · 1997
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2003
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2003
Cited alongside, same era.
Online geometric optimization in the bandit setting against an adaptive adversary
H. Brendan McMahan and Avrim Blum · 2004
Cited alongside, same era.
Online convex optimization in the bandit setting: gradient descent without a gradient
Abraham D. Flaxman, Adam Tauman Kalai, and H. Brendan McMahan · 2005
Cited alongside, same era.
Efficient algorithms for online decision problems
Adam Kalai and Santosh Vempala · 2005
Cited alongside, same era.
On following the perturbed leader in the bandit setting
Jussi Kujala and Tapio Elomaa · 2005
Cited alongside, same era.
Prediction, Learning, and Games
Nicolò Cesa-Bianchi and Gábor Lugosi · 2006
Cited alongside, same era.
Multi-armed bandit allocation indices
John Gittins, Kevin Glazebrook, and Richard Weber · 2011
Later among the works it cites.
Interior-point methods for full-information and bandit online learning
Jacob Abernethy, Elad Hazan, and Alexander Rakhlin · 2012
Later among the works it cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Later among the works it cites.
Reliability Engineering
E.A. Elsayed · 2012
Later among the works it cites.
Hyperparameter tuning in bandit-based adaptive operator selection
Maciej Pacula, Jason Ansel, Saman Amarasinghe, and Una-May O’Reilly · 2012
Later among the works it cites.
Relax and randomize: From value to algorithms
Sasha Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robbing the bandit: less regret in online geometric optimization against an adaptive adversary
V. Dani and T. P. Hayes · 2006
Cited alongside, same era.
The price of bandit information for online optimization
Varsha Dani, Thomas Hayes, and Sham Kakade · 2008
Cited alongside, same era.
Minimax policies for adversarial and stochastic bandits
Jean-Yves Audibert and Sébastien Bubeck · 2009
Cited alongside, same era.
Monte-carlo tree search in poker using expected reward distributions
Guy Van den Broeck, Kurt Driessens, and Jan Ramon · 2009
Cited alongside, same era.
Minimax policies for combinatorial prediction games
Jean-Yves Audibert, Sébastien Bubeck, and Gábor Lugosi · 2011
Cited alongside, same era.
Later among the works it cites.
Prediction by random-walk perturbation
Luc Devroye, Gábor Lugosi, and Gergely Neu · 2013
Later among the works it cites.
An efficient algorithm for learning with semi-bandit feedback
Gergely Neu and Gábor Bartók · 2013
Later among the works it cites.
Online linear optimization via smoothing
Jacob Abernethy, Chansoo Lee, Abhinav Sinha, and Ambuj Tewari · 2014
Later among the works it cites.
Efficient learning by implicit exploration in bandit problems with side observations
Tomáš Kocák, Gergely Neu, Michal Valko, and Remi Munos · 2014
Later among the works it cites.
Follow the leader with dropout perturbations
Tim Van Erven, Wojciech Kotlowski, and Manfred K Warmuth · 2014
Later among the works it cites.