Fetching the paper…
Reading the bibliography…
We prove a new minimax theorem connecting the worst-case Bayesian regret and minimax regret under partial monitoring with no assumptions on the space of signals or decisions of the adversary.
Extensive games and the problem of information, contributions to the theory of games II
H.W. Kuhn · 1953
Earlier work this paper cites.
On general minimax theorems
M. Sion · 1958
Earlier work this paper cites.
Information and strategies in dynamic games
P. Bernhard · 1992
Earlier work this paper cites.
Minimizing regret: The general case
A. Rustichini · 1999
Earlier work this paper cites.
Nonstationary zero sum stochastic games with incomplete observation
D. Leao, J. B. R. do Val, and M. D. Fragoso · 2000
Earlier work this paper cites.
Foundations of modern probability
O. Kallenberg · 2002
Earlier work this paper cites.
On-line learning with imperfect monitoring
S. Mannor and N. Shimkin · 2003
Earlier work this paper cites.
Zero-sum stochastic games with partial information
M. K. Ghosh, D. McDonald, and S. Sinha · 2004
Earlier work this paper cites.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Regret minimization under partial monitoring
N. Cesa-Bianchi, G. Lugosi, and G. Stoltz · 2006
Earlier work this paper cites.
Measure theory , volume 2
V. I. Bogachev · 2007
Earlier work this paper cites.
Introduction to nonparametric estimation
A. B. Tsybakov · 2008
Earlier work this paper cites.
A stochastic view of optimal regret through minimax duality
J. Abernethy, A. Agarwal, P. L. Bartlett, and A. Rakhlin · 2009
Earlier work this paper cites.
Minimax regret of finite partial-monitoring games in stochastic environments
G. Bartók, D. Pál, and Cs. Szepesvári · 2011
Cited alongside, same era.
Approachability of convex sets in games with partial monitoring
V. Perchet · 2011
Cited alongside, same era.
An adaptive algorithm for finite stochastic partial monitoring
G. Bartók, N. Zolghadr, and Cs. Szepesvári · 2012
Cited alongside, same era.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Towards minimax policies for online linear optimization with bandit feedback
S. Bubeck, N. Cesa-Bianchi, and S. Kakade · 2012
Cited alongside, same era.
No internal regret via neighborhood watch
D. Foster and A. Rakhlin · 2012
Cited alongside, same era.
Efficient partial monitoring with prior information
H. P. Vanchinathan, G. Bartók, and A. Krause · 2014
Later among the works it cites.
Online learning with feedback graphs: Beyond bandits
N. Alon, N. Cesa-Bianchi, O. Dekel, and T. Koren · 2015
Later among the works it cites.
Bandit convex optimization: T \sqrt{T} regret in one dimension
S. Bubeck, O. Dekel, T. Koren, and Y. Peres · 2015
Later among the works it cites.
Regret lower bound and optimal algorithm in finite stochastic partial monitoring
J. Komiyama, J. Honda, and H. Nakagawa · 2015
Later among the works it cites.
Towards optimal algorithms for prediction with expert advice
N. Gravin, Y. Peres, and B. Sivan · 2016
Later among the works it cites.
An information-theoretic analysis of Thompson sampling
D. Russo and B. Van Roy · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thompson sampling: An asymptotically optimal finite-time analysis
E. Kaufmann, N. Korda, and R. Munos · 2012
Cited alongside, same era.
Further optimal regret bounds for Thompson sampling
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
Toward a classification of finite partial-monitoring games
A. Antos, G. Bartók, D. Pál, and Cs. Szepesvári · 2013
Cited alongside, same era.
Online linear optimization via smoothing
J. D. Abernethy, C. Lee, A. Sinha, and A. Tewari · 2014
Cited alongside, same era.
Partial monitoring—classification, regret bounds, and algorithms
G. Bartók, D. P. Foster, D. Pál, A. Rakhlin, and Cs. Szepesvári · 2014
Cited alongside, same era.
Set-valued approachability and online learning with partial monitoring
S. Mannor, V. Perchet, and G. Stoltz · 2014
Cited alongside, same era.
Fundamentals of nonparametric Bayesian inference , volume 44
S. Ghosal and A. van der Vaart · 2017
Later among the works it cites.
Learning to optimize via information-directed sampling
D. Russo and B. Van Roy · 2017
Later among the works it cites.
Sparsity, variance and curvature in multi-armed bandits
S. Bubeck, M. Cohen, and Y. Li · 2018
Later among the works it cites.
More adaptive algorithms for adversarial bandits
C-Y. Wei and H. Luo · 2018
Later among the works it cites.
Cleaning up the neighbourhood: A full classification for adversarial partial monitoring
T. Lattimore and Cs. Szepesvári · 2019
Closest in time.
Bandit Algorithms
T. Lattimore and Cs. Szepesvári · 2019
Closest in time.