Fetching the paper…
Reading the bibliography…
We propose the first contextual bandit algorithm that is parameter-free, efficient, and optimal in terms of dynamic regret.
Tracking the best expert
Mark Herbster and Manfred K Warmuth · 1998
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Tracking a small set of experts by mixing past posteriors
Olivier Bousquet and Manfred K Warmuth · 2002
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2008
Earlier work this paper cites.
Adapting to a changing environment: the brownian restless bandits
Aleksandrs Slivkins and Eli Upfal · 2008
Earlier work this paper cites.
Efficient learning algorithms for changing environments
Elad Hazan and Comandur Seshadhri · 2009
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert E Schapire · 2011
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
M. Dudík, D. Hsu, S. Kale, N. Karampatziakis, J. Langford, L. Reyzin, and T. Zhang · 2011
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
Aurélien Garivier and Eric Moulines · 2011
Earlier work this paper cites.
Putting bayes to sleep
Dmitry Adamskiy, Manfred K Warmuth, and Wouter M Koolen · 2012
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert E Schapire · 2014
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Cited alongside, same era.
Non-stationary stochastic optimization
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2015
Cited alongside, same era.
Online optimization: Competing with dynamic comparators
Ali Jadbabaie, Alexander Rakhlin, Shahin Shahrampour, and Karthik Sridharan · 2015
Cited alongside, same era.
Achieving All with No Parameters: AdaNormalHedge
Haipeng Luo and Robert E. Schapire · 2015
Cited alongside, same era.
Tracking the best expert in non-stationary stochastic environments
Chen-Yu Wei, Yi-Te Hong, and Chi-Jen Lu · 2016
Later among the works it cites.
Tracking slowly moving clairvoyant: optimal dynamic regret of online learning with true and noisy gradient
Tianbao Yang, Lijun Zhang, Rong Jin, and Jinfeng Yi · 2016
Later among the works it cites.
Online learning for changing environments using coin betting
Kwang-Sung Jun, Francesco Orabona, Stephen Wright, Rebecca Willett, et al · 2017
Later among the works it cites.
Improved dynamic regret for non-degenerate functions
Lijun Zhang, Tianbao Yang, Jinfeng Yi, Jing Rong, and Zhi-Hua Zhou · 2017
Later among the works it cites.
Adaptively tracking the best arm with an unknown number of distribution changes
Peter Auer, Pratik Gajane, and Ronald Ortner · 2018
Later among the works it cites.
A change-detection based framework for piecewise-stationary multi-armed bandit problem
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The computational power of optimization in online learning
Elad Hazan and Tomer Koren · 2016
Cited alongside, same era.
Multi-armed bandits: Competing with optimal sequences
Zohar S Karnin and Oren Anava · 2016
Cited alongside, same era.
Bistro: An efficient relaxation-based method for contextual bandits
Alexander Rakhlin and Karthik Sridharan · 2016
Cited alongside, same era.
Efficient algorithms for adversarial contextual learning
Vasilis Syrgkanis, Akshay Krishnamurthy, and Robert E Schapire
Cited in the paper.
Improved regret bounds for oracle-based adversarial contextual bandits
Vasilis Syrgkanis, Haipeng Luo, Akshay Krishnamurthy, and Robert E Schapire
Cited in the paper.
Fang Liu, Joohyun Lee, and Ness Shroff · 2018
Later among the works it cites.
Efficient contextual bandits in non-stationary worlds
Haipeng Luo, Chen-Yu Wei, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Dynamic regret of strongly adaptive methods
Lijun Zhang, Tianbao Yang, Rong Jin, and Zhi-Hua Zhou · 2018
Later among the works it cites.
Learning to optimize under non-stationarity
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2019
Closest in time.