Fetching the paper…
Reading the bibliography…
We develop a novel and generic algorithm for the adversarial multi-armed bandit problem (or more generally the combinatorial semi-bandit problem).
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Yoav Freund and Robert E Schapire · 1999
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Hannan consistency in on-line learning in case of unbounded losses under partial monitoring
Chamy Allenberg, Peter Auer, Laszlo Gyorfi, and György Ottucsák · 2006
Earlier work this paper cites.
Competing in the dark: An efficient algorithm for bandit linear optimization
Jacob D Abernethy, Elad Hazan, and Alexander Rakhlin · 2008
Earlier work this paper cites.
The kl-ucb algorithm for bounded stochastic bandits and beyond
Aurélien Garivier and Olivier Cappé · 2011
Earlier work this paper cites.
The best of both worlds: stochastic and adversarial bandits
Sébastien Bubeck and Aleksandrs Slivkins · 2012
Earlier work this paper cites.
Online optimization with gradual variations
Chao-Kai Chiang, Tianbao Yang, Chia-Jung Lee, Mehrdad Mahdavi, Chi-Jen Lu, Rong Jin, and Shenghuo Zhu · 2012
Earlier work this paper cites.
Regret in online combinatorial optimization
Jean-Yves Audibert, Sébastien Bubeck, and Gábor Lugosi · 2013
Earlier work this paper cites.
Beating bandits in gradually evolving worlds
Chao-Kai Chiang, Chia-Jung Lee, and Chi-Jen Lu · 2013
Earlier work this paper cites.
Follow the leader if you can, hedge if you must
Steven De Rooij, Tim Van Erven, Peter D Grünwald, and Wouter M Koolen · 2014
Earlier work this paper cites.
A second-order bound with excess losses
Pierre Gaillard, Gilles Stoltz, and Tim Van Erven · 2014
Earlier work this paper cites.
One practical algorithm for both stochastic and adversarial bandits
Yevgeny Seldin and Aleksandrs Slivkins · 2014
Cited alongside, same era.
Adaptivity and optimism: An improved exponentiated gradient algorithm
Jacob Steinhardt and Percy Liang · 2014
Cited alongside, same era.
Near-optimal no-regret algorithms for zero-sum games
Constantinos Daskalakis, Alan Deckelbaum, and Anthony Kim · 2015
Cited alongside, same era.
Second-order quantile methods for experts and combinatorial games
Wouter M Koolen and Tim Van Erven · 2015
Cited alongside, same era.
Optimally confident ucb: Improved regret for finite-armed bandits
Tor Lattimore · 2015
Cited alongside, same era.
Achieving all with no parameters: Adanormalhedge
Haipeng Luo and Robert E Schapire · 2015
Learning in games: Robustness of fast convergence
Dylan J Foster, Zhiyuan Li, Thodoris Lykouris, Karthik Sridharan, and Eva Tardos · 2016
Later among the works it cites.
Refined lower bounds for adversarial bandits
Sébastien Gerchinovitz and Tor Lattimore · 2016
Later among the works it cites.
Introduction to online convex optimization
Elad Hazan et al · 2016
Later among the works it cites.
Coin betting and parameter-free online learning
Francesco Orabona and Dávid Pál · 2016
Later among the works it cites.
Metagrad: Multiple learning rates in online learning
Tim van Erven and Wouter M Koolen · 2016
Later among the works it cites.
Corralling a band of bandit algorithms
Alekh Agarwal, Haipeng Luo, Behnam Neyshabur, and Robert E Schapire · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
First-order regret bounds for combinatorial semi-bandits
Gergely Neu · 2015
Cited alongside, same era.
Fast convergence of regularized learning in games
Vasilis Syrgkanis, Alekh Agarwal, Haipeng Luo, and Robert E Schapire · 2015
Cited alongside, same era.
An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits
Peter Auer and Chao-Kai Chiang · 2016
Cited alongside, same era.
Kernel-based methods for bandit convex optimization
Sébastien Bubeck, Ronen Eldan, and Yin Tat Lee · 2016
Cited alongside, same era.
Anytime optimal algorithms in stochastic multi-armed bandits
Rémy Degenne and Vianney Perchet · 2016
Cited alongside, same era.
Better algorithms for benign bandits
Elad Hazan and Satyen Kale
Cited in the paper.
Sparsity, variance and curvature in multi-armed bandits
Sébastien Bubeck, Michael B. Cohen, and Yuanzhi Li · 2017
Later among the works it cites.
Online learning without prior information
Ashok Cutkosky and Kwabena Boahen · 2017
Later among the works it cites.
Small-loss bounds for online learning with partial information
Thodoris Lykouris, Karthik Sridharan, and Eva Tardos · 2017
Later among the works it cites.
A survey of algorithms and analysis for adaptive online learning
H Brendan McMahan · 2017
Later among the works it cites.
An improved parametrization and analysis of the exp3++ algorithm for stochastic and adversarial bandits
Yevgeny Seldin and Gábor Lugosi · 2017
Later among the works it cites.