Fetching the paper…
Reading the bibliography…
We present a new algorithm for the contextual bandit learning problem, where the learner repeatedly takes one of $K$ actions in response to the observed context, and observes the reward only for that chosen action.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Predicting nearly as well as the best pruning of a decision tree
David P. Helmbold and Robert E. Schapire · 1997
Earlier work this paper cites.
Reinforcement learning, an introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Beating the holdout: Bounds for k-fold and progressive cross-validation
Avrim Blum, Adam Kalai, and John Langford · 1999
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2002
Earlier work this paper cites.
Rcv1: A new benchmark collection for text categorization research
David D Lewis, Yiming Yang, Tony G Rose, and Fan Li · 2004
Cited alongside, same era.
The epoch-greedy algorithm for contextual multi-armed bandits
John Langford and Tong Zhang · 2007
Cited alongside, same era.
The offset tree for learning with partial labels
Alina Beygelzimer and John Langford · 2009
Cited alongside, same era.
Tighter bounds for multi-armed bandits with expert advice
H. Brendan McMahan and Matthew Streeter · 2009
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Cited alongside, same era.
Efficient optimal learning for contextual bandits
Miroslav Dudík, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang
Cited in the paper.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Later among the works it cites.
An empirical evaluation of Thompson sampling
Olivier Chapelle and Lihong Li · 2011
Later among the works it cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Later among the works it cites.
Generalized Thompson sampling for contextual bandits
Lihong Li · 2013
Later among the works it cites.
Interactive machine learning, January 2014
John Langford · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li
Cited in the paper.