Fetching the paper…
Reading the bibliography…
This work addresses the problem of regret minimization in non-stochastic multi-armed bandit problems, focusing on performance guarantees that hold with high probability.
Approximation to Bayes risk in repeated play
J. Hannan · 1957
Earlier work this paper cites.
On tail probabilities for martingales
D. A. Freedman · 1975
Earlier work this paper cites.
V. Vovk · 1990
Earlier work this paper cites.
The weighted majority algorithm
N. Littlestone and M. Warmuth · 1994
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R. E. Schapire · 1997
Earlier work this paper cites.
Tracking the best expert
M. Herbster and M. Warmuth · 1998
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 2002
Earlier work this paper cites.
Efficient algorithms for online decision problems
A. Kalai and S. Vempala · 2005
Earlier work this paper cites.
Prediction, Learning, and Games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
High-probability regret bounds for bandit online linear optimization
P. L. Bartlett, V. Dani, T. P. Hayes, S. Kakade, A. Rakhlin, and A. Tewari · 2008
Cited alongside, same era.
Minimax policies for adversarial and stochastic bandits
J.-Y. Audibert and S. Bubeck · 2009
Cited alongside, same era.
Tighter bounds for multi-armed bandits with expert advice
H. B. McMahan and M. Streeter · 2009
Cited alongside, same era.
Regret bounds and minimax policies under partial monitoring
J.-Y. Audibert and S. Bubeck · 2010
Cited alongside, same era.
Contextual bandit algorithms with supervised learning guarantees
A. Beygelzimer, J. Langford, L. Li, L. Reyzin, and R. E. Schapire · 2011
Cited alongside, same era.
Better algorithms for benign bandits
E. Hazan and S. Kale · 2011
Cited alongside, same era.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
S. Bubeck and N. Cesa-Bianchi · 2012
Later among the works it cites.
Mirror descent meets fixed share (and feels no regret)
N. Cesa-Bianchi, P. Gaillard, G. Lugosi, and G. Stoltz · 2012
Later among the works it cites.
PAC-Bayes-Bernstein inequality for martingales and its application to multiarmed bandits
Y. Seldin, N. Cesa-Bianchi, P. Auer, F. Laviolette, and J. Shawe-Taylor · 2012
Later among the works it cites.
Online learning with predictable sequences
A. Rakhlin and K. Sridharan · 2013
Later among the works it cites.
Nonstochastic multi-armed bandits with graph-structured feedback
N. Alon, N. Cesa-Bianchi, C. Gentile, S. Mannor, Y. Mansour, and O. Shamir · 2014
Later among the works it cites.
Volumetric spanners: an efficient exploration basis for learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
From Bandits to Experts: On the Value of Side-Observations
S. Mannor and O. Shamir · 2011
Cited alongside, same era.
From Bandits to Experts: A Tale of Domination and Independence
N. Alon, N. Cesa-Bianchi, C. Gentile, and Y. Mansour · 2012
Cited alongside, same era.
Towards minimax policies for online linear optimization with bandit feedback
S. Bubeck, N. Cesa-Bianchi, and S. M. Kakade · 2012
Cited alongside, same era.
E. Hazan, Z. Karnin, and R. Meka · 2014
Later among the works it cites.
Efficient learning by implicit exploration in bandit problems with side observations
T. Kocák, G. Neu, M. Valko, and R. Munos · 2014
Later among the works it cites.
First-order regret bounds for combinatorial semi-bandits
G. Neu · 2015
Closest in time.