Fetching the paper…
Reading the bibliography…
The principle of optimism in the face of uncertainty is one of the most widely used and successful ideas in multi-armed bandits and reinforcement learning.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Adaptive routing with end-to-end feedback: Distributed learning and geometric approaches
Baruch Awerbuch and Robert D Kleinberg · 2004
Earlier work this paper cites.
Nearly tight bounds for the continuum-armed bandit problem
Robert D Kleinberg · 2005
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
High-probability regret bounds for bandit online linear optimization
Peter L Bartlett, Varsha Dani, Thomas Hayes, Sham Kakade, Alexander Rakhlin, and Ambuj Tewari · 2008
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
On upper-confidence bound policies for non-stationary bandit problems
Aurélien Garivier and Eric Moulines · 2008
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham M Kakade, and Matthias Seeger · 2009
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappe, Aurélien Garivier, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Nonparametric bandits with covariates
Philippe Rigollet and Assaf Zeevi · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Cited alongside, same era.
Efficient optimal learning for contextual bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Cited alongside, same era.
Contextual bandit learning with predictable rewards
Alekh Agarwal, Miroslav Dudík, Satyen Kale, John Langford, and Robert Schapire · 2012
Cited alongside, same era.
Large-scale bandit problems and kwik learning
Jacob Abernethy, Kareem Amin, Michael Kearns, and Moez Draief · 2013
Cited alongside, same era.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Alberto Bietti, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Practical contextual bandits with regression oracles
Dylan J Foster, Alekh Agarwal, Miroslav Dudík, Haipeng Luo, and Robert E Schapire · 2018
Later among the works it cites.
Semi-parametric efficient policy learning with continuous actions
Victor Chernozhukov, Mert Demirer, Greg Lewis, and Vasilis Syrgkanis · 2019
Later among the works it cites.
Active learning for cost-sensitive classification
Akshay Krishnamurthy, Alekh Agarwal, Tzu-Kuo Huang, Hal Daumé III, and John Langford · 2019
Later among the works it cites.
Contextual bandits with continuous actions: Smoothing, zooming, and adapting
Akshay Krishnamurthy, John Langford, Aleksandrs Slivkins, and Chicheng Zhang · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Cited alongside, same era.
Contextual bandits with similarity information
Aleksandrs Slivkins · 2014
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Provably optimal algorithms for generalized linear contextual bandits
Lihong Li, Yu Lu, and Dengyong Zhou · 2017
Cited alongside, same era.
Later among the works it cites.
Introduction to multi-armed bandits
Aleksandrs Slivkins et al · 2019
Later among the works it cites.
Adapting to misspecification in contextual bandits
Dylan J Foster, Claudio Gentile, Mehryar Mohri, and Julian Zimmert · 2020
Closest in time.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
Dylan J Foster and Alexander Rakhlin · 2020
Closest in time.
David Simchi-Levi and Yunzong Xu · 2020
Closest in time.
Counterfactual optimism: Rate optimal regret for stochastic contextual mdps
Orin Levy, Asaf Cassel, Alon Cohen, and Yishay Mansour · 2022
Closest in time.
Optimism in face of a context: Regret guarantees for stochastic contextual mdp
Orin Levy and Yishay Mansour · 2023
Closest in time.