Fetching the paper…
Reading the bibliography…
We study the power of different types of adaptive (nonoblivious) adversaries in the setting of prediction with expert advice, under both full-information and bandit feedback.
Asymptotically efficient adaptive allocation rules for the multiarmed bandit problem with switching cost
R. Agrawal, M.V. Hedge, and D. Teneketzis · 1988
Earlier work this paper cites.
The weighted majority algorithm
N. Littlestone and M.K. Warmuth · 1994
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R.E. Schapire · 1997
Earlier work this paper cites.
Online computation and competitive analysis
A. Borodin and R. El-Yaniv · 1998
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. Schapire · 2002
Earlier work this paper cites.
Sequential strategies for loss functions with memory
N. Merhav, E. Ordentlich, G. Seroussi, and M.J. Weinberger · 2002
Earlier work this paper cites.
A survey on the bandit problem with switching costs
T. Jun · 2004
Cited alongside, same era.
Online geometric optimization in the bandit setting against an adaptive adversary
H. B. McMahan and A. Blum · 2004
Cited alongside, same era.
Efficient algorithms for online decision problems
A. Kalai and S. Vempala · 2005
Cited alongside, same era.
Online learning with delayed label feedback
C. Mesterharm · 2005
Cited alongside, same era.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Cited alongside, same era.
Robbing the bandit: Less regret in online geometric optimization against an adaptive adversary
V. Dani and T. P. Hayes · 2006
Cited alongside, same era.
Improved second-order bounds for prediction with expert advice
N. Cesa-Bianchi, Y. Mansour, and G. Stoltz · 2007
Later among the works it cites.
Adaptive bandits: Towards the best history-dependent strategy
O. Maillard and R. Munos · 2010
Later among the works it cites.
Online regret bounds for Markov decision processes with deterministic transitions
R. Ortner · 2010
Later among the works it cites.
Online bandit learning against an adaptive adversary: from regret to policy regret
R. Arora, O. Dekel, and A. Tewari · 2012
Later among the works it cites.
Regret minimization for reserve prices in second-price auctions
N. Cesa-Bianchi, C. Gentile, and Y. Mansour · 2013
Closest in time.
On the complexity of bandit and derivative-free stochastic convex optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O. Shamir · 2013
Closest in time.