Fetching the paper…
Reading the bibliography…
We propose a new online algorithm for cumulative regret minimization in a stochastic linear bandit.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Approximation to Bayes risk in repeated play
James Hannan · 1957
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Non-Uniform Random Variate Generation
Luc Devroye · 1986
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto · 1998
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Efficient algorithms for online decision problems
Adam Kalai and Santosh Vempala · 2005
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappe, Aurelien Garivier, and Csaba Szepesvari · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Paat Rusmevichientong and John Tsitsiklis · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, David Pal, and Csaba Szepesvari · 2011
Cited alongside, same era.
An empirical evaluation of Thompson sampling
Olivier Chapelle and Lihong Li · 2011
Cited alongside, same era.
An efficient algorithm for learning with semi-bandit feedback
Gergely Neu and Gabor Bartok · 2013
Cited alongside, same era.
Sub-sampling for multi-armed bandits
Akram Baransi, Odalric-Ambrym Maillard, and Shie Mannor · 2014
Cited alongside, same era.
Thompson sampling with the online bootstrap
Dean Eckles and Maurits Kaptein · 2014
Cited alongside, same era.
Thompson sampling for complex online problems
Aditya Gopalan, Shie Mannor, and Yishay Mansour · 2014
Cited alongside, same era.
A practical method for solving contextual bandit problems using decision trees
Adam Elmachtoub, Ryan McNellis, Sechan Oh, and Marek Petrik · 2017
Later among the works it cites.
Scalable generalized linear bandits: Online computation and hashing
Kwang-Sung Jun, Aniruddha Bhargava, Robert Nowak, and Rebecca Willett · 2017
Later among the works it cites.
Provably optimal algorithms for generalized linear contextual bandits
Lihong Li, Yu Lu, and Dengyong Zhou · 2017
Later among the works it cites.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Later among the works it cites.
BBQ-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems
Zachary Lipton, Xiujun Li, Jianfeng Gao, Lihong Li, Faisal Ahmed, and Li Deng · 2018
Later among the works it cites.
Customized nonlinear bandits for online response selection in neural conversation models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spectral bandits for smooth graph functions
Michal Valko, Remi Munos, Branislav Kveton, and Tomas Kocak · 2014
Cited alongside, same era.
Efficient Thompson sampling for online matrix-factorization recommendation
Jaya Kawale, Hung Bui, Branislav Kveton, Long Tran-Thanh, and Sanjay Chawla · 2015
Cited alongside, same era.
Bootstrapped Thompson sampling and deep exploration
Ian Osband and Benjamin Van Roy · 2015
Cited alongside, same era.
Personalized recommendation via parameter-free contextual bandits
Liang Tang, Yexi Jiang, Lei Li, Chunqiu Zeng, and Tao Li · 2015
Cited alongside, same era.
Further optimal regret bounds for Thompson sampling
Shipra Agrawal and Navin Goyal
Cited in the paper.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal
Cited in the paper.
Bing Liu, Tong Yu, Ian Lane, and Ole Mengshoel · 2018
Later among the works it cites.
Deep Bayesian bandits showdown: An empirical comparison of Bayesian deep networks for Thompson sampling
Carlos Riquelme, George Tucker, and Jasper Snoek · 2018
Later among the works it cites.
A tutorial on Thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2018
Later among the works it cites.
New insights into bootstrapping for bandits
Sharan Vaswani, Branislav Kveton, Zheng Wen, Anup Rao, Mark Schmidt, and Yasin Abbasi-Yadkori · 2018
Later among the works it cites.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvari · 2019
Closest in time.