Fetching the paper…
Reading the bibliography…
We initiate the study of learning in contextual bandits with the help of loss predictors.
Semiparametric efficiency in multivariate regression models with missing data
James M Robins and Andrea Rotnitzky · 1995
Earlier work this paper cites.
Using and combining predictors that specialize
Yoav Freund, Robert E Schapire, Yoram Singer, and Manfred K Warmuth · 1997
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2008
Earlier work this paper cites.
Extracting certainty from uncertainty: Regret bounded by variation in costs
Elad Hazan and Satyen Kale · 2010
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Earlier work this paper cites.
Better algorithms for benign bandits
Elad Hazan and Satyen Kale · 2011
Earlier work this paper cites.
Equipping experts/bandits with long-term memory
Kai Zheng, Haipeng Luo, Ilias Diakonikolas, and Liwei Wang · 2011
Earlier work this paper cites.
Challenging the empirical mean and empirical variance: a deviation study
Olivier Catoni · 2012
Cited alongside, same era.
Online optimization with gradual variations
Chao-Kai Chiang, Tianbao Yang, Chia-Jung Lee, Mehrdad Mahdavi, Chi-Jen Lu, Rong Jin, and Shenghuo Zhu · 2012
Cited alongside, same era.
Beating bandits in gradually evolving worlds
Chao-Kai Chiang, Chia-Jung Lee, and Chi-Jen Lu · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Cited alongside, same era.
Doubly robust policy evaluation and optimization
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li · 2014
Cited alongside, same era.
Adaptivity and optimism: An improved exponentiated gradient algorithm
Jacob Steinhardt and Percy Liang · 2014
Tracking the best expert in non-stationary stochastic environments
Chen-Yu Wei, Yi-Te Hong, and Chi-Jen Lu · 2016
Later among the works it cites.
Lecture notes 21 of introduction to online learning
Haipeng Luo · 2017
Later among the works it cites.
Make the minority great again: First-order regret bound for contextual bandits
Zeyuan Allen-Zhu, Sébastien Bubeck, and Yuanzhi Li · 2018
Later among the works it cites.
More adaptive algorithms for adversarial bandits
Chen-Yu Wei and Haipeng Luo · 2018
Later among the works it cites.
Improved path-length regret bounds for bandits
Sébastien Bubeck, Yuanzhi Li, Haipeng Luo, and Chen-Yu Wei · 2019
Later among the works it cites.
A new algorithm for non-stationary contextual bandits: Efficient, optimal and parameter-free
Yifang Chen, Chung-Wei Lee, Haipeng Luo, and Chen-Yu Wei · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Strongly adaptive online learning
Amit Daniely, Alon Gonen, and Shai Shalev-Shwartz · 2015
Cited alongside, same era.
Fast convergence of regularized learning in games
Vasilis Syrgkanis, Alekh Agarwal, Haipeng Luo, and Robert E Schapire · 2015
Cited alongside, same era.
The computational power of optimization in online learning
Elad Hazan and Tomer Koren · 2016
Cited alongside, same era.
Online learning with predictable sequences
Alexander Rakhlin and Karthik Sridharan
Cited in the paper.
Optimization, learning, and games with predictable sequences
Sasha Rakhlin and Karthik Sridharan
Cited in the paper.
Efficient algorithms for adversarial contextual learning
Vasilis Syrgkanis, Akshay Krishnamurthy, and Robert E Schapire
Cited in the paper.
Later among the works it cites.
Contextual bandits with continuous actions: Smoothing, zooming, and adapting
Akshay Krishnamurthy, John Langford, Aleksandrs Slivkins, and Chicheng Zhang · 2019
Later among the works it cites.
Mean estimation and regression under heavy-tailed distributions: A survey
Gábor Lugosi and Shahar Mendelson · 2019
Later among the works it cites.