Fetching the paper…
Reading the bibliography…
This paper investigates stochastic and adversarial combinatorial multi-armed bandit problems.
On cliques in graphs
J.W. Moon and L. Moser · 1965
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Herbert Robbins · 1985
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Asymptotically efficient allocation rules for the multiarmed bandit problem with multiple plays-part i: iid rewards
Venkatachalam Anantharam, Pravin Varaiya, and Jean Walrand · 1987
Earlier work this paper cites.
A constructive proof of the representation theorem for polyhedral sets based on fundamental definitions
H. D. Sherali · 1987
Earlier work this paper cites.
Asymptotically efficient adaptive choice of control laws in controlled markov chains
Todd L. Graves and Tze Leung Lai · 1997
Earlier work this paper cites.
Finite time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Combinatorial Optimization: Polyhedra and Efficiency
Alexander Schrijver · 2003
Earlier work this paper cites.
Information theory and statistics: A tutorial
I. Csiszár and P.C. Shields · 2004
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Cited alongside, same era.
Prediction, learning, and games
Nicolò Cesa-Bianchi and Gábor Lugosi · 2006
Cited alongside, same era.
The on-line shortest path problem under partial monitoring
András György, Tamás Linder, Gábor Lugosi, and György Ottucsák · 2007
Cited alongside, same era.
Learning permutations with exponential weights
David P. Helmbold and Manfred K. Warmuth · 2009
Cited alongside, same era.
Learning multiuser channel allocations in cognitive radio networks: A combinatorial multi-armed bandit formulation
Yi Gai, Bhaskar Krishnamachari, and Rahul Jain · 2010
Cited alongside, same era.
Non-stochastic bandit slate problems
Satyen Kale, Lev Reyzin, and Robert Schapire · 2010
Cited alongside, same era.
Towards minimax policies for online linear optimization with bandit feedback
Sébastien Bubeck, Nicolò Cesa-Bianchi, and Sham M. Kakade · 2012
Later among the works it cites.
Combinatorial multi-armed bandit: General framework and applications
Wei Chen, Yajun Wang, and Yang Yuan · 2013
Later among the works it cites.
Regret in online combinatorial optimization
Jean-Yves Audibert, Sébastien Bubeck, and Gábor Lugosi · 2013
Later among the works it cites.
Matroid bandits: Fast combinatorial optimization with learning
Branislav Kveton, Zheng Wen, Azin Ashkan, Hoda Eydgahi, and Brian Eriksson · 2014
Later among the works it cites.
Bandit online optimization over the permutahedron
Nir Ailon, Kohei Hatano, and Eiji Takimoto · 2014
Later among the works it cites.
Lipschitz bandits: Regret lower bounds and optimal algorithms
Stefan Magureanu, Richard Combes, and Alexandre Proutiere · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The KL-UCB algorithm for bounded stochastic bandits and beyond
Aurélien Garivier and Olivier Cappé · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Cited alongside, same era.
Combinatorial bandits
Nicolò Cesa-Bianchi and Gábor Lugosi · 2012
Cited alongside, same era.
Combinatorial network optimization with unknown variables: Multi-armed bandits with linear rewards and individual observations
Yi Gai, Bhaskar Krishnamachari, and Rahul Jain · 2012
Cited alongside, same era.
Later among the works it cites.
Unimodal bandits: Regret lower bounds and optimal algorithms
Richard Combes and Alexandre Proutiere · 2014
Later among the works it cites.
Tight regret bounds for stochastic combinatorial semi-bandits
Branislav Kveton, Zheng Wen, Azin Ashkan, and Csaba Szepesvari · 2015
Closest in time.
Efficient learning in large-scale combinatorial semi-bandits
Zheng Wen, Azin Ashkan, Hoda Eydgahi, and Branislav Kveton · 2015
Closest in time.
First-order regret bounds for combinatorial semi-bandits
Gergely Neu · 2015
Closest in time.