Fetching the paper…
Reading the bibliography…
We address the problem of learning in an online, bandit setting where the learner must repeatedly select among $K$ actions, but only receives partial feedback based on its choices.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
On tail probabilities for martingales
David A. Freedman · 1975
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Associative reinforcement learning: Functions in k k -DNF
Leslie Pack Kaelbling · 1994
Earlier work this paper cites.
A generalization of Sauer’s lemma
David Haussler and Philip M. Long · 1995
Earlier work this paper cites.
Predicting nearly as well as the best pruning of a decision tree
David P. Helmbold and Robert E. Schapire · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2002
Cited alongside, same era.
From batch to transductive online learning
Sham M. Kakade and Adam Kalai · 2005
Cited alongside, same era.
Prediction, Learning, and Games
Nicolò Cesa-Bianchi and Gabor Lugosi · 2006
Cited alongside, same era.
Experience-efficient learning in associative bandit problems
Alexander L. Strehl, Chris Mesterharm, Michael L. Littman, and Haym Hirsh · 2006
Cited alongside, same era.
The epoch-greedy algorithm for contextual multi-armed bandits
John Langford and Tong Zhang · 2007
Cited alongside, same era.
Online models for content optimization
Deepak Agarwal, Bee-Chung Chen, Pradheep Elango, Nitin Motgi, Seung-Taek Park, Raghu Ramakrishnan, Scott Roy, and Joe Zachariah · 2008
Cited alongside, same era.
High-probability regret bounds for bandit online linear optimization
Peter Bartlett, Varsha Dani, Thomas Hayes, Sham Kakade, Alexander Rakhlin, and Ambuj Tewari · 2008
Later among the works it cites.
Efficient bandit algorithms for online multiclass prediction
Sham M. Kakade, Shai Shalev-Shwartz, and Ambuj Tewari · 2008
Later among the works it cites.
Agnostic online learning
Shai Ben-david, Dávid Pál, and Shai Shalev-shwartz · 2009
Later among the works it cites.
Hybrid stochastic-adversarial on-line learning
Alessandro Lazaric and Rémi Munos · 2009
Later among the works it cites.
Tighter bounds for multi-armed bandits with expert advice
Brendan McMahan and Matthew Streeter · 2009
Later among the works it cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…