Fetching the paper…
Reading the bibliography…
Most contextual bandit algorithms minimize regret against the best fixed policy, a questionable benchmark for non-stationary environments that are ubiquitous in applications.
Bistro: An efficient relaxation-based method for contextual bandits
Alexander Rakhlin and Karthik Sridharan · 1985
Earlier work this paper cites.
Tracking the best expert
Mark Herbster and Manfred K Warmuth · 1998
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Tracking a small set of experts by mixing past posteriors
Olivier Bousquet and Manfred K Warmuth · 2002
Earlier work this paper cites.
Adaptive algorithms for online decision problems
Elad Hazan and C. Seshadhri · 2007
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2008
Earlier work this paper cites.
Mortal multi-armed bandits
Deepayan Chakrabarti, Ravi Kumar, Filip Radlinski, and Eli Upfal · 2009
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Cited alongside, same era.
Strongly adaptive regret implies optimally dynamic regret
Lijun Zhang, Tianbao Yang, Rong Jin, and Zhi-Hua Zhou · 2011
Cited alongside, same era.
The best of both worlds: Stochastic and adversarial bandits
Sébastien Bubeck and Aleksandrs Slivkins · 2012
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Cited alongside, same era.
Non-stationary stochastic optimization
An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits
Peter Auer and Chao-Kai Chiang · 2016
Later among the works it cites.
The computational power of optimization in online learning
Elad Hazan and Tomer Koren · 2016
Later among the works it cites.
Multi-armed bandits: Competing with optimal sequences
Zohar S Karnin and Oren Anava · 2016
Later among the works it cites.
Tracking the best expert in non-stationary stochastic environments
Chen-Yu Wei, Yi-Te Hong, and Chi-Jen Lu · 2016
Later among the works it cites.
Corralling a band of bandit algorithms
Alekh Agarwal, Haipeng Luo, Behnam Neyshabur, and Robert E Schapire · 2017
Closest in time.
More adaptive algorithms for adversarial bandits
Chen-Yu Wei and Haipeng Luo · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2015
Cited alongside, same era.
Strongly adaptive online learning
Amit Daniely, Alon Gonen, and Shai Shalev-Shwartz · 2015
Cited alongside, same era.
Bistro: An efficient relaxation-based method for contextual bandits
Alexander Rakhlin and Karthik Sridharan
Cited in the paper.
Efficient algorithms for adversarial contextual learning
Vasilis Syrgkanis, Akshay Krishnamurthy, and Robert E Schapire
Cited in the paper.
Improved regret bounds for oracle-based adversarial contextual bandits
Vasilis Syrgkanis, Haipeng Luo, Akshay Krishnamurthy, and Robert E Schapire
Cited in the paper.
Closest in time.