Fetching the paper…
Reading the bibliography…
We address online combinatorial optimization when the player has a prior over the adversary's sequence of losses.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. Thompson · 1933
Earlier work this paper cites.
On tail probabilities for martingales
David A Freedman · 1975
Earlier work this paper cites.
How to use expert advice
Nicolo Cesa-Bianchi, Yoav Freund, David Haussler, David P Helmbold, Robert E Schapire, and Manfred K Warmuth · 1997
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Hannan consistency in on-line learning in case of unbounded losses under partial monitoring
C. Allenberg, P. Auer, L. Györfi, and G. Ottucsák · 2006
Earlier work this paper cites.
Hedging structured concepts
W. Koolen, M. Warmuth, and J. Kivinen · 2010
Earlier work this paper cites.
An Empirical Evaluation of Thompson Sampling
O. Chapelle and L. Li · 2011
Earlier work this paper cites.
From bandits to experts: On the value of side-observations
Shie Mannor and Ohad Shamir · 2011
Earlier work this paper cites.
Analysis of Thompson sampling for the multi-armed bandit problem
S. Agrawal and N. Goyal · 2012
Cited alongside, same era.
Brownian motion and stochastic calculus
Ioannis Karatzas and Steven Shreve · 2012
Cited alongside, same era.
Regret in online combinatorial optimization
J.Y. Audibert, S. Bubeck, and G. Lugosi · 2014
Cited alongside, same era.
Online learning with feedback graphs: Beyond bandits
Noga Alon, Nicolo Cesa-Bianchi, Ofer Dekel, and Tomer Koren · 2015
Cited alongside, same era.
Bandit convex optimization: T \sqrt{T} regret in one dimension
S. Bubeck, O. Dekel, T. Koren, and Y. Peres · 2015
Cited alongside, same era.
An information-theoretic analysis of thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Cited alongside, same era.
Make the minority great again: First-order regret bound for contextual bandits
Zeyuan Allen-Zhu, Sebastien Bubeck, and Yuanzhi Li · 2018
Later among the works it cites.
Small-loss bounds for online learning with partial information
T. Lykouris, K. Sridharan, and E. Tardos · 2018
Later among the works it cites.
Analysis of Thompson Sampling for Graphical Bandits Without the Graphs
Fang Liu, Zizhan Zheng, and Ness Shroff · 2018
Later among the works it cites.
An information-theoretic approach to minimax regret in partial monitoring
Tor Lattimore and Csaba Szepesvári · 2019
Closest in time.
Connections Between Mirror Descent, Thompson Sampling and the Information Ratio
Julian Zimmert and Tor Lattimore · 2019
Closest in time.
Feedback graph regret bounds for Thompson Sampling and UCB
Thodoris Lykouris, Eva Tardos, and Drishti Wali · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Open problem: First-order regret bounds for contextual bandits
Alekh Agarwal, Akshay Krishnamurthy, John Langford, Haipeng Luo, et al · 2017
Cited alongside, same era.
Thompson sampling for stochastic bandits with graph feedback
Aristide CY Tossou, Christos Dimitrakakis, and Devdatt P Dubhashi · 2017
Cited alongside, same era.
Closest in time.
Efficient first-order contextual bandits: Prediction, allocation, and triangular discrimination
Dylan J Foster and Akshay Krishnamurthy · 2021
Closest in time.
Mirror descent and the information ratio
Tor Lattimore and Andras Gyorgy · 2021
Closest in time.