Fetching the paper…
Reading the bibliography…
The information-theoretic analysis by Russo and Van Roy (2014) in combination with minimax duality has proved a powerful tool for the analysis of online learning algorithms in full and partial information settings.
Efficient methods for large-scale convex optimization problems
A. S. Nemirovsky · 1979
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. S. Nemirovsky and D. B. Yudin · 1983
Earlier work this paper cites.
Minimizing regret: The general case
A. Rustichini · 1999
Earlier work this paper cites.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Regret minimization under partial monitoring
N. Cesa-Bianchi, G. Lugosi, and G. Stoltz · 2006
Earlier work this paper cites.
Competing in the dark: An efficient algorithm for bandit linear optimization
J. D. Abernethy, E. Hazan, and A. Rakhlin · 2008
Earlier work this paper cites.
A stochastic view of optimal regret through minimax duality
J. Abernethy, A. Agarwal, P. L. Bartlett, and A. Rakhlin · 2009
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
J.-Y. Audibert and S. Bubeck · 2009
Earlier work this paper cites.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
S. Bubeck and N. Cesa-Bianchi · 2012
Earlier work this paper cites.
No internal regret via neighborhood watch
D. Foster and A. Rakhlin · 2012
Cited alongside, same era.
Toward a classification of finite partial-monitoring games
A. Antos, G. Bartók, D. Pál, and Cs. Szepesvári · 2013
Cited alongside, same era.
Partial monitoring—classification, regret bounds, and algorithms
G. Bartók, D. P. Foster, D. Pál, A. Rakhlin, and Cs. Szepesvári · 2014
Cited alongside, same era.
Online learning with feedback graphs: Beyond bandits
N. Alon, N. Cesa-Bianchi, O. Dekel, and T. Koren · 2015
Cited alongside, same era.
Bandit convex optimization: T \sqrt{T} regret in one dimension
S. Bubeck, O. Dekel, T. Koren, and Y. Peres · 2015
Cited alongside, same era.
Online learning with feedback graphs without the graphs
A. Cohen, T. Hazan, and T. Koren · 2016
Cited alongside, same era.
Bandits on graphs and structures, 2016
M. Valko · 2016
Later among the works it cites.
Nonstochastic multi-armed bandits with graph-structured feedback
N. Alon, N. Cesa-Bianchi, C. Gentile, S. Mannor, Y. Mansour, and O. Shamir · 2017
Later among the works it cites.
Sparsity, variance and curvature in multi-armed bandits
S. Bubeck, M. Cohen, and Y. Li · 2018
Later among the works it cites.
An information-theoretic analysis for Thompson sampling with many actions
S. Dong and B. Van Roy · 2018
Later among the works it cites.
First-order regret analysis of thompson sampling
S. Bubeck and M. Sellke · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards optimal algorithms for prediction with expert advice
N. Gravin, Y. Peres, and B. Sivan · 2016
Cited alongside, same era.
Introduction to online convex optimization
E. Hazan · 2016
Cited alongside, same era.
An information-theoretic analysis of Thompson sampling
D. Russo and B. Van Roy · 2016
Cited alongside, same era.
Learning to optimize via information-directed sampling
D. Russo and B. Van Roy
Cited in the paper.
Learning to optimize via posterior sampling
D. Russo and B. Van Roy
Cited in the paper.
Cleaning up the neighbourhood: A full classification for adversarial partial monitoring
T. Lattimore and Cs. Szepesvári · 2019
Closest in time.
Bandit Algorithms
T. Lattimore and Cs. Szepesvári · 2019
Closest in time.
An information-theoretic approach to minimax regret in partial monitoring
T. Lattimore and Cs. Szepesvári · 2019
Closest in time.