Fetching the paper…
Reading the bibliography…
Partial monitoring is a rich framework for sequential decision making under uncertainty that generalizes many well known bandit models, including linear, combinatorial and dueling bandits.
Estimation des densités: risque minimax
J Bretagnolle and C Huber · 1979
Earlier work this paper cites.
A one-armed bandit problem with a concomitant variable
Michael Woodroofe · 1979
Earlier work this paper cites.
Associative reinforcement learning using linear probabilistic concepts
Naoki Abe and Philip M. Long · 1999
Earlier work this paper cites.
Minimizing regret: The general case
A. Rustichini · 1999
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2003
Earlier work this paper cites.
Regret minimization under partial monitoring
N. Cesa-Bianchi, G. Lugosi, and G. Stoltz · 2006
Earlier work this paper cites.
Stochastic Linear Optimization under Bandit Feedback
Varsha Dani, Thomas P. Hayes, and Sham M. Kakade · 2008
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
J. Langford and T. Zhang · 2008
Earlier work this paper cites.
Interactively optimizing information retrieval systems as a dueling bandits problem
Yisong Yue and Thorsten Joachims · 2009
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham M Kakade, and Matthias Seeger · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Minimax regret of finite partial-monitoring games in stochastic environments
G. Bartók, D. Pál, and Cs. Szepesvári · 2011
Earlier work this paper cites.
Online Learning for Linearly Parametrized Control Problems
Yasin Abbasi-Yadkori · 2012
Earlier work this paper cites.
An adaptive algorithm for finite stochastic partial monitoring
G. Bartók, N. Zolghadr, and Cs. Szepesvári · 2012
Earlier work this paper cites.
Partial monitoring with side information
Gábor Bartók and Csaba Szepesvári · 2012
Cited alongside, same era.
Combinatorial bandits
Nicolo Cesa-Bianchi and Gábor Lugosi · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Toward a classification of finite partial-monitoring games
A. Antos, G. Bartók, D. Pál, and Cs. Szepesvári · 2013
Cited alongside, same era.
Partial monitoring—classification, regret bounds, and algorithms
G. Bartók, D. P. Foster, D. Pál, A. Rakhlin, and Cs. Szepesvári · 2014
Cited alongside, same era.
Combinatorial partial monitoring game with linear feedback and its applications
Tian Lin, Bruno Abrahao, Robert Kleinberg, John Lui, and Wei Chen · 2014
Cited alongside, same era.
On kernelized multi-armed bandits
Sayak Ray Chowdhury and Aditya Gopalan · 2017
Later among the works it cites.
Preferential bayesian optimization
Javier González, Zhenwen Dai, Andreas Damianou, and Neil D Lawrence · 2017
Later among the works it cites.
Following the leader and fast rates in online linear prediction: Curved constraint sets and other regularities
R. Huang, T. Lattimore, A. György, and Cs. Szepesvári · 2017
Later among the works it cites.
Sparsity, variance and curvature in multi-armed bandits
S. Bubeck, M. Cohen, and Y. Li · 2018
Later among the works it cites.
Gaussian processes and kernel methods: A review on connections and equivalences
Motonobu Kanagawa, Philipp Hennig, Dino Sejdinovic, and Bharath K Sriperumbudur · 2018
Later among the works it cites.
Information directed sampling and bandits with heteroscedastic noise
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to optimize via information-directed sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Efficient partial monitoring with prior information
Hastagiri Vanchinathan, Gábor Bartók, and Andreas Krause · 2014
Cited alongside, same era.
Utility-based dueling bandits as a partial monitoring game
Pratik Gajane and Tanguy Urvoy · 2015
Cited alongside, same era.
Regret lower bound and optimal algorithm in finite stochastic partial monitoring
J. Komiyama, J. Honda, and H. Nakagawa · 2015
Cited alongside, same era.
Phased exploration with greedy exploitation in stochastic combinatorial partial monitoring games
Sougata Chaudhuri and Ambuj Tewari · 2016
Cited alongside, same era.
Refined lower bounds for adversarial bandits
S. Gerchinovitz and T. Lattimore · 2016
Cited alongside, same era.
Johannes Kirschner and Andreas Krause · 2018
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2018
Later among the works it cites.
Advancements in dueling bandits
Yanan Sui, Masrour Zoghi, Katja Hofmann, and Yisong Yue · 2018
Later among the works it cites.
Sequential experimental design for transductive linear bandits
Tanner Fiez, Lalit Jain, Kevin G Jamieson, and Lillian Ratliff · 2019
Later among the works it cites.
Adaptive exploration in linear contextual bandit
Botao Hao, Tor Lattimore, and Csaba Szepesvari · 2019
Later among the works it cites.
Adaptive and safe bayesian optimization in high dimensions via one-dimensional subspaces
Johannes Kirschner, Mojmir Mutny, Nicole Hiller, Rasmus Ischebeck, and Andreas Krause · 2019
Later among the works it cites.
Cleaning up the neighborhood: A full classification for adversarial partial monitoring
Tor Lattimore and Csaba Szepesvári · 2019
Later among the works it cites.
Exploration by optimisation in partial monitoring
Tor Lattimore and Csaba Szepesvari · 2019
Later among the works it cites.
An information-theoretic approach to minimax regret in partial monitoring
Tor Lattimore and Csaba Szepesvári · 2019
Later among the works it cites.