Fetching the paper…
Reading the bibliography…
A fundamental challenge in contextual bandits is to develop flexible, general-purpose algorithms with computational requirements no worse than classical supervised learning tasks such as classification and regression.
A game of prediction with expert advice
Vladimir Vovk · 1995
Earlier work this paper cites.
Scale-sensitive dimensions, uniform convergence, and learnability
Noga Alon, Shai Ben-David, Nicolo Cesa-Bianchi, and David Haussler · 1997
Earlier work this paper cites.
Prediction, learning, uniform convergence, and scale-sensitive dimensions
Peter L Bartlett and Philip M Long · 1998
Earlier work this paper cites.
Competitive on-line linear regression
Vladimir Vovk · 1998
Earlier work this paper cites.
Associative reinforcement learning using linear probabilistic concepts
Naoki Abe and Philip M Long · 1999
Earlier work this paper cites.
Information-theoretic determination of minimax rates of convergence
Yuhong Yang and Andrew Barron · 1999
Earlier work this paper cites.
Relative loss bounds for on-line density estimation with the exponential family of distributions
Katy S Azoury and Manfred K Warmuth · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Naoki Abe, Alan W Biermann, and Philip M Long · 2003
Earlier work this paper cites.
Entropy and the combinatorial dimension
Shahar Mendelson and Roman Vershynin · 2003
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gabor Lugosi · 2006
Earlier work this paper cites.
Metric entropy in competitive on-line prediction
Vladimir Vovk · 2006
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2008
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Alexandre B Tsybakov · 2008
Earlier work this paper cites.
Cryptographic hardness for learning intersections of halfspaces
Adam R Klivans and Alexander A Sherstov · 2009
Earlier work this paper cites.
Nonparametric bandits with covariates
Philippe Rigollet and Assaf Zeevi · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E Schapire · 2011
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Cited alongside, same era.
Efficient learning of generalized linear and single index models with isotonic regression
Sham M Kakade, Varun Kanade, Ohad Shamir, and Adam Kalai · 2011
Cited alongside, same era.
Knows what it knows: a framework for self-aware learning
Lihong Li, Michael L Littman, Thomas J Walsh, and Alexander L Strehl · 2011
Cited alongside, same era.
Contextual bandits with similarity information
Aleksandrs Slivkins · 2011
Cited alongside, same era.
On the universality of online mirror descent
Nathan Srebro, Karthik Sridharan, and Ambuj Tewari · 2011
Cited alongside, same era.
Online-to-confidence-set conversions and application to sparse stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvari · 2012
Contextual semibandits via supervised learning oracles
Akshay Krishnamurthy, Alekh Agarwal, and Miro Dudik · 2016
Later among the works it cites.
BISTRO: An efficient relaxation-based method for contextual bandits
Alexander Rakhlin and Karthik Sridharan · 2016
Later among the works it cites.
Open problem: First-order regret bounds for contextual bandits
Alekh Agarwal, Akshay Krishnamurthy, John Langford, and Haipeng Luo · 2017
Later among the works it cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Later among the works it cites.
Provably optimal algorithms for generalized linear contextual bandits
Lihong Li, Yu Lu, and Dengyong Zhou · 2017
Later among the works it cites.
Empirical entropy, minimax regret and minimax risk
Alexander Rakhlin, Karthik Sridharan, and Alexandre B Tsybakov · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Contextual bandit learning with predictable rewards
Alekh Agarwal, Miroslav Dudík, Satyen Kale, John Langford, and Robert E Schapire · 2012
Cited alongside, same era.
Statistical learning and sequential prediction, 2012
Alexander Rakhlin and Karthik Sridharan · 2012
Cited alongside, same era.
Large-scale bandit problems and KWIK learning
Jacob Abernethy, Kareem Amin, Michael Kearns, and Moez Draief · 2013
Cited alongside, same era.
Sparsity regret bounds for individual sequences in online linear regression
Sébastien Gerchinovitz · 2013
Cited alongside, same era.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Cited alongside, same era.
Finite-time analysis of kernelised contextual bandits
Michal Valko, Nathan Korda, Rémi Munos, Ilias Flaounas, and Nello Cristianini · 2013
Cited alongside, same era.
Later among the works it cites.
From ads to interventions: Contextual bandits in mobile health
Ambuj Tewari and Susan A Murphy · 2017
Later among the works it cites.
Make the minority great again: First-order regret bound for contextual bandits
Zeyuan Allen-Zhu, Sébastien Bubeck, and Yuanzhi Li · 2018
Later among the works it cites.
Alberto Bietti, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Contextual bandits with surrogate losses: Margin bounds and efficient algorithms
Dylan J Foster and Akshay Krishnamurthy · 2018
Later among the works it cites.
Practical contextual bandits with regression oracles
Dylan J Foster, Alekh Agarwal, Miroslav Dudík, Haipeng Luo, and Robert E Schapire · 2018
Later among the works it cites.
New insights into bootstrapping for bandits
Sharan Vaswani, Branislav Kveton, Zheng Wen, Anup Rao, Mark Schmidt, and Yasin Abbasi-Yadkori · 2018
Later among the works it cites.
Model selection for contextual bandits
Dylan J Foster, Akshay Krishnamurthy, and Haipeng Luo · 2019
Later among the works it cites.
Bandits and experts in metric spaces
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal · 2019
Later among the works it cites.
Garbage in, reward out: Bootstrapping exploration in multi-armed bandits
Branislav Kveton, Csaba Szepesvari, Sharan Vaswani, Zheng Wen, Tor Lattimore, and Mohammad Ghavamzadeh · 2019
Later among the works it cites.
Learning with good feature representations in bandits and in rl with a generative model
Tor Lattimore and Csaba Szepesvari · 2019
Later among the works it cites.
Nearly minimax-optimal regret for linearly parameterized bandits
Yingkai Li, Yining Wang, and Yuan Zhou · 2019
Later among the works it cites.
Comments on the Du-Kakade-Wang-Yang lower bounds
Benjamin Van Roy and Shi Dong · 2019
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S Du, Sham M Kakade, Ruosong Wang, and Lin F Yang · 2020
Closest in time.
Efficient and robust algorithms for adversarial linear contextual bandits
Gergely Neu and Julia Olkhovskaya · 2020
Closest in time.