Fetching the paper…
Reading the bibliography…
We propose an online algorithm for cumulative regret minimization in a stochastic multi-armed bandit.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Popoviciu’s inequality on variances
Tiberiu Popoviciu · 1935
Earlier work this paper cites.
Approximation to Bayes risk in repeated play
James Hannan · 1957
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Efficient algorithms for online decision problems
Adam Kalai and Santosh Vempala · 2005
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
Learning diverse rankings with multi-armed bandits
Filip Radlinski, Robert Kleinberg, and Thorsten Joachims · 2008
Earlier work this paper cites.
A dynamic Bayesian network click model for web search ranking
Olivier Chapelle and Ya Zhang · 2009
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappe, Aurelien Garivier, and Csaba Szepesvari · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, David Pal, and Csaba Szepesvari · 2011
Earlier work this paper cites.
The KL-UCB algorithm for bounded stochastic bandits and beyond
Aurelien Garivier and Olivier Cappe · 2011
Cited alongside, same era.
Further optimal regret bounds for Thompson sampling
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
An efficient algorithm for learning with semi-bandit feedback
Gergely Neu and Gabor Bartok · 2013
Cited alongside, same era.
Thompson sampling for complex online problems
Aditya Gopalan, Shie Mannor, and Yishay Mansour · 2014
Cited alongside, same era.
Efficient Thompson sampling for online matrix-factorization recommendation
Jaya Kawale, Hung Bui, Branislav Kveton, Long Tran-Thanh, and Sanjay Chawla · 2015
Cited alongside, same era.
Scalable generalized linear bandits: Online computation and hashing
Kwang-Sung Jun, Aniruddha Bhargava, Robert Nowak, and Rebecca Willett · 2017
Later among the works it cites.
Provably optimal algorithms for generalized linear contextual bandits
Lihong Li, Yu Lu, and Dengyong Zhou · 2017
Later among the works it cites.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Later among the works it cites.
BBQ-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems
Zachary Lipton, Xiujun Li, Jianfeng Gao, Lihong Li, Faisal Ahmed, and Li Deng · 2018
Later among the works it cites.
Customized nonlinear bandits for online response selection in neural conversation models
Bing Liu, Tong Yu, Ian Lane, and Ole Mengshoel · 2018
Later among the works it cites.
Deep Bayesian bandits showdown: An empirical comparison of Bayesian deep networks for Thompson sampling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cascading bandits: Learning to rank in the cascade model
Branislav Kveton, Csaba Szepesvari, Zheng Wen, and Azin Ashkan · 2015
Cited alongside, same era.
DCM bandits: Learning to rank with multiple clicks
Sumeet Katariya, Branislav Kveton, Csaba Szepesvari, and Zheng Wen · 2016
Cited alongside, same era.
Online stochastic linear optimization under one-bit feedback
Lijun Zhang, Tianbao Yang, Rong Jin, Yichi Xiao, and Zhi-Hua Zhou · 2016
Cited alongside, same era.
Linear Thompson sampling revisited
Marc Abeille and Alessandro Lazaric · 2017
Cited alongside, same era.
Carlos Riquelme, George Tucker, and Jasper Snoek · 2018
Later among the works it cites.
A tutorial on Thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2018
Later among the works it cites.
Perturbed-history exploration in stochastic linear bandits
Branislav Kveton, Csaba Szepesvari, Mohammad Ghavamzadeh, and Craig Boutilier · 2019
Closest in time.
Garbage in, reward out: Bootstrapping exploration in multi-armed bandits
Branislav Kveton, Csaba Szepesvari, Sharan Vaswani, Zheng Wen, Mohammad Ghavamzadeh, and Tor Lattimore · 2019
Closest in time.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvari · 2019
Closest in time.