Fetching the paper…
Reading the bibliography…
We derive an algorithm that achieves the optimal (within constants) pseudo-regret in both adversarial and stochastic multi-armed bandits without prior knowledge of the regime and time horizon.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Possible generalization of Boltzmann-Gibbs statistics
Constantino Tsallis · 1988
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolò Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
Jean-Yves Audibert and Sébastien Bubeck · 2009
Earlier work this paper cites.
Regret bounds and minimax policies under partial monitoring
Jean-Yves Audibert and Sébastien Bubeck · 2010
Earlier work this paper cites.
Bandits games and clustering foundations
Sébastien Bubeck · 2010
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Earlier work this paper cites.
The best of both worlds: Stochastic and adversarial bandits
Sébastien Bubeck and Aleksandrs Slivkins · 2012
Earlier work this paper cites.
Thompson sampling: An optimal finite time analysis
Emilie Kaufmann, Nathaniel Korda, and Rémi Munos · 2012
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz · 2012
Cited alongside, same era.
Kullback-Leibler upper confidence bounds for optimal sequential allocation
Olivier Cappé, Aurélien Garivier, Odalric-Ambrym Maillard, Rémi Munos, and Gilles Stoltz · 2013
Cited alongside, same era.
Online linear optimization via smoothing
Jacob Abernethy, Chansoo Lee, Abhinav Sinha, and Ambuj Tewari · 2014
Cited alongside, same era.
Reducing dueling bandits to cardinal bandits
Nir Ailon, Zohar Karnin, and Thorsten Joachims · 2014
Cited alongside, same era.
One practical algorithm for both stochastic and adversarial bandits
Yevgeny Seldin and Aleksandrs Slivkins · 2014
Cited alongside, same era.
Corralling a band of bandit algorithms
Alekh Agarwal, Haipeng Luo, Behnam Neyshabur, and Robert E Schapire · 2017
Later among the works it cites.
An improved parametrization and analysis of the EXP3++ algorithm for stochastic and adversarial bandits
Yevgeny Seldin and Gábor Lugosi · 2017
Later among the works it cites.
Best of both worlds: Stochastic & adversarial best-arm identification
Yasin Abbasi-Yadkori, Peter Bartlett, Victor Gabillon, Alan Malek, and Michal Valko · 2018
Closest in time.
What doubling tricks can and can’t do for multi-armed bandits
Lilian Besson and Emilie Kaufmann · 2018
Closest in time.
Stochastic bandits robust to adversarial corruptions
Thodoris Lykouris, Vahab Mirrokni, and Renato Paes Leme · 2018
Closest in time.
More adaptive algorithms for adversarial bandits
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fighting bandits with a new kind of smoothness
Jacob D Abernethy, Chansoo Lee, and Ambuj Tewari · 2015
Cited alongside, same era.
A generalized online mirror descent with applications to classification and regression
Francesco Orabona, Koby Crammer, and Nicolò Cesa-Bianchi · 2015
Cited alongside, same era.
Convex analysis
Ralph Tyrell Rockafellar · 2015
Cited alongside, same era.
An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits
Peter Auer and Chao-Kai Chiang · 2016
Cited alongside, same era.
Learning in games: Robustness of fast convergence
Dylan J Foster, Zhiyuan Li, Thodoris Lykouris, Karthik Sridharan, and Eva Tardos · 2016
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer
Cited in the paper.
Chen-Yu Wei and Haipeng Luo · 2018
Closest in time.
Better algorithms for stochastic bandits with adversarial corruptions
Anupam Gupta, Tomer Koren, and Kunal Talwar · 2019
Closest in time.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2019
Closest in time.
Connections between mirror descent, Thompson sampling, and the information ratio
Julian Zimmert and Tor Lattimore · 2019
Closest in time.
An optimal algorithm for stochastic and adversarial bandits
Julian Zimmert and Yevgeny Seldin · 2019
Closest in time.
Beating stochastic and adversarial semi-bandits optimally and simultaneously
Julian Zimmert, Haipeng Luo, and Chen-Yu Wei · 2019
Closest in time.