Fetching the paper…
Reading the bibliography…
We present and study a partial-information model of online learning, where a decision maker repeatedly chooses from a finite set of actions, and observes some subset of the associated losses.
On tail probabilities for martingales
D.A. Freedman · 1975
Earlier work this paper cites.
New results on the independence number
Y. Caro · 1979
Earlier work this paper cites.
A greedy heuristic for the set-covering problem
V. Chvatal · 1979
Earlier work this paper cites.
A lower bound on the stability number of a simple graph
V. K. Wey · 1981
Earlier work this paper cites.
On the independence number of random graphs
A. M. Frieze · 1990
Earlier work this paper cites.
Aggregating strategies
V. G. Vovk · 1990
Earlier work this paper cites.
The weighted majority algorithm
Nick Littlestone and Manfred K. Warmuth · 1994
Earlier work this paper cites.
How to use expert advice
N. Cesa-Bianchi, Y. Freund, D. Haussler, D. P. Helmbold, R. E. Schapire, and M. K. Warmuth · 1997
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E. Schapire · 1997
Cited alongside, same era.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2002
Cited alongside, same era.
The probabilistic method
N. Alon and J. H. Spencer · 2004
Cited alongside, same era.
Improved second-order bounds for prediction with expert advice
Nicolò Cesa-Bianchi, Yishay Mansour, and Gilles Stoltz · 2005
Cited alongside, same era.
Efficient algorithms for online decision problems
A. Kalai and S. Vempala · 2005
Cited alongside, same era.
Information Incomplète et Regret Interne en Prédiction de Suites Individuelles
Gilles Stoltz · 2005
Cited alongside, same era.
How social relationships affect user similarities
Alan Said, Ernesto W De Luca, and Sahin Albayrak · 2010
Later among the works it cites.
From bandits to experts: On the value of side-observations
S. Mannor and O. Shamir · 2011
Later among the works it cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Later among the works it cites.
Leveraging side observations in stochastic bandits
Stéphane Caron, Branislav Kveton, Marc Lelarge, and Smriti Bhagat · 2012
Later among the works it cites.
Combinatorial bandits
Nicolò Cesa-Bianchi and Gábor Lugosi · 2012
Later among the works it cites.
Online bandit learning against an adaptive adversary: from regret to policy regret
Ofer Dekel, Ambuj Tewari, and Raman Arora · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Cited alongside, same era.
Minimax policies for adversarial and stochastic bandits
Jean-Yves Audibert and Sébastien Bubeck · 2009
Cited alongside, same era.
From bandits to experts: A tale of domination and independence
N. Alon, N. Cesa-Bianchi, C. Gentile, and Y. Mansour · 2013
Later among the works it cites.
Efficient learning by implicit exploration in bandit problems with side observations
T. Kocàk, G. Neu, M. Valko, and R. Munos · 2014
Closest in time.