Fetching the paper…
Reading the bibliography…
We provide the first oracle efficient sublinear regret algorithms for adversarial versions of the contextual bandit problem.
A generalization of sampling without replacement from a finite universe
Horvitz, Daniel G and Thompson, Donovan J · 1952
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit pproblem
Auer, Peter, Cesa-Bianchi, Nicolo, Freund, Yoav, and Schapire, Robert E · 1995
Earlier work this paper cites.
Characterizations of learnability for classes of (0,…, n)-valued functions
Ben-David, Shai, Cesa-Bianchi, Nicolo, Haussler, David, and Long, Philip M · 1995
Earlier work this paper cites.
A generalization of sauer’s lemma
Haussler, David and Long, Philip M · 1995
Earlier work this paper cites.
Online learning versus offline learning
Ben-David, Shai, Kushilevitz, Eyal, and Mansour, Yishay · 1997
Earlier work this paper cites.
How to use expert advice
Cesa-Bianchi, Nicolo, Freund, Yoav, Haussler, David, Helmbold, David P, Schapire, Robert E, and Warmuth, Manfred K · 1997
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Yoav and Schapire, Robert E · 1997
Earlier work this paper cites.
Tracking the best expert
Herbster, Mark and Warmuth, Manfred K · 1998
Earlier work this paper cites.
Adaptive online prediction by following the perturbed leader
Hutter, Marcus and Poland, Jan · 2005
Cited alongside, same era.
From batch to transductive online learning
Kakade, Sham M and Kalai, Adam · 2005
Cited alongside, same era.
Efficient algorithms for online decision problems
Kalai, Adam and Vempala, Santosh · 2005
Cited alongside, same era.
Online linear optimization and adaptive routing
Awerbuch, Baruch and Kleinberg, Robert · 2008
Cited alongside, same era.
The epoch-greedy algorithm for multi-armed bandits with side information
Langford, John and Zhang, Tong · 2008
Cited alongside, same era.
Extracting certainty from uncertainty: regret bounded by variation in costs
Hazan, Elad and Kale, Satyen · 2010
Cited alongside, same era.
Online submodular minimization for combinatorial structures
Jegelka, Stefanie and Bilmes, Jeff A · 2011
Later among the works it cites.
Online optimization with gradual variations
Chiang, Chao-Kai, Yang, Tianbao, Lee, Chia-Jung, Mahdavi, Mehrdad, Lu, Chi-Jen, Jin, Rong, and Zhu, Shenghuo · 2012
Later among the works it cites.
Online submodular minimization
Hazan, Elad and Kale, Satyen · 2012
Later among the works it cites.
An efficient algorithm for learning with semi-bandit feedback
Neu, Gergely and Bartók, Gábor · 2013
Later among the works it cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Agarwal, Alekh, Hsu, Daniel, Kale, Satyen, Langford, John, Li, Lihong, and Schapire, Robert E · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient online learning via randomized rounding
Cesa-Bianchi, Nicolo and Shamir, Ohad · 2011
Cited alongside, same era.
Efficient optimal learning for contextual bandits
Dudík, Miroslav, Hsu, Daniel, Kale, Satyen, Karampatziakis, Nikos, Langford, John, Reyzin, Lev, and Zhang, Tong · 2011
Cited alongside, same era.
Optimization, learning, and games with predictable sequences
Rakhlin, Alexander and Sridharan, Karthik
Cited in the paper.
Online learning with predictable sequences
Rakhlin, Alexander and Sridharan, Karthik
Cited in the paper.
Daskalakis, Constantinos and Syrgkanis, Vasilis · 2015
Later among the works it cites.
Achieving all with no parameters: Adanormalhedge
Luo, Haipeng and Schapire, Robert E · 2015
Later among the works it cites.
Fast convergence of regularized learning in games
Syrgkanis, Vasilis, Agarwal, Alekh, Luo, Haipeng, and Schapire, Robert E · 2015
Later among the works it cites.