Fetching the paper…
Reading the bibliography…
We consider the adversarial multi-armed bandit problem under delayed feedback.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2002
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
Jean-Yves Audibert and Sébastien Bubeck · 2009
Earlier work this paper cites.
Online markov decision processes under bandit feedback
Gergely Neu, András György, Csaba Szepesvári, and András Antos · 2010
Earlier work this paper cites.
Online learning under delayed feedback
Pooria Joulani, András György, and Csaba Szepesvári · 2013
Earlier work this paper cites.
Online Markov decision processes under bandit feedback
Gergely Neu, András. György, Csaba Szepesvári, and András Antos · 2014
Earlier work this paper cites.
Online learning with adversarial delays
Kent Quanrud and Daniel Khashabi · 2015
Cited alongside, same era.
Delay-tolerant online convex optimization: Unified analysis and adaptive-gradient algorithms
Pooria Joulani, András György, and Csaba Szepesvári · 2016
Cited alongside, same era.
A survey of algorithms and analysis for adaptive online learning
H. Brendan McMahan · 2017
Cited alongside, same era.
Online EXP3 learning in adversarial bandits with delayed feedback
Ilai Bistritz, Zhengyuan Zhou, Xi Chen, Nicholas Bambos, and Jose Blanchet · 2019
Cited alongside, same era.
First-order regret bounds for combinatorial semi-bandits
Gergely Neu
Cited in the paper.
Explore no more: Improved high-probability regret bounds for non-stochastic bandits
Gergely Neu
Cited in the paper.
Delay and cooperation in nonstochastic bandits
Nicolò Cesa-Bianchi, Claudio Gentile, and Yishay Mansour · 2019
Later among the works it cites.
Nonstochastic multiarmed bandits with unrestricted delays
Tobias Sommer Thune, Nicolò Cesa-Bianchi, and Yevgeny Seldin · 2019
Later among the works it cites.
An optimal algorithm for adversarial bandits with arbitrary delays
Julian Zimmert and Yevgeny Seldin · 2019
Later among the works it cites.
A modular analysis of adaptive (non-)convex optimization: Optimism, composite objectives, variance reduction, and variational bounds
Pooria Joulani, András György, and Csaba Szepesvári · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…