Fetching the paper…
Reading the bibliography…
We address the problem of learning in an online setting where the learner repeatedly observes features, selects among a set of actions, and receives reward for the action taken.
On general minimax theorems
Maurice Sion · 1958
Earlier work this paper cites.
On tail probabilities for martingales
David A. Freedman · 1975
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham Kakade, and Matthias Seeger · 1975
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R. E. Schapire · 1997
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Cited alongside, same era.
From batch to transductive online learning
Sham M. Kakade and Adam Kalai · 2005
Cited alongside, same era.
Efficient algorithms for online decision problems
Adam Tauman Kalai and Santosh Vempala · 2005
Cited alongside, same era.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2006
Cited alongside, same era.
Adaptive online gradient descent
P. L. Bartlett, E. Hazan, and A. Rakhlin · 2007
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer
Cited in the paper.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire
Cited in the paper.
The epoch-greedy algorithm for contextual multi-armed bandits
John Langford and Tong Zhang · 2007
Later among the works it cites.
Error correcting tournaments
Alina Beygelzimer, John Langford, and Pradeep Ravikumar · 2009
Later among the works it cites.
Slow learners are fast
J. Langford, A. Smola, and M. Zinkevich · 2009
Later among the works it cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…