Fetching the paper…
Reading the bibliography…
We propose a bandit algorithm that explores by randomizing its history of rewards.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Stochastic Processes
Joseph Doob · 1953
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Bootstrap methods for standard errors, confidence intervals, and other measures of statistical accuracy
Bradley Efron and Robert Tibshirani · 1986
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto · 1998
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappe, Aurelien Garivier, and Csaba Szepesvari · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert Schapire · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Paat Rusmevichientong and John Tsitsiklis · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, David Pal, and Csaba Szepesvari · 2011
Cited alongside, same era.
The KL-UCB algorithm for bounded stochastic bandits and beyond
Aurelien Garivier and Olivier Cappe · 2011
Cited alongside, same era.
Scikit-learn: Machine learning in Python
Fabian Pedregosa, Gael Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake VanderPlas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Edouard Duchesnay · 2011
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Cited alongside, same era.
Sub-sampling for multi-armed bandits
Akram Baransi, Odalric-Ambrym Maillard, and Shie Mannor · 2014
Cited alongside, same era.
Personalized recommendation via parameter-free contextual bandits
Liang Tang, Yexi Jiang, Lei Li, Chunqiu Zeng, and Tao Li · 2015
Later among the works it cites.
Online stochastic linear optimization under one-bit feedback
Lijun Zhang, Tianbao Yang, Rong Jin, Yichi Xiao, and Zhi-Hua Zhou · 2016
Later among the works it cites.
A practical method for solving contextual bandit problems using decision trees
Adam Elmachtoub, Ryan McNellis, Sechan Oh, and Marek Petrik · 2017
Later among the works it cites.
Scalable generalized linear bandits: Online computation and hashing
Kwang-Sung Jun, Aniruddha Bhargava, Robert Nowak, and Rebecca Willett · 2017
Later among the works it cites.
Provably optimal algorithms for generalized linear contextual bandits
Lihong Li, Yu Lu, and Dengyong Zhou · 2017
Later among the works it cites.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dean Eckles and Maurits Kaptein · 2014
Cited alongside, same era.
Thompson sampling for complex online problems
Aditya Gopalan, Shie Mannor, and Yishay Mansour · 2014
Cited alongside, same era.
Efficient Thompson sampling for online matrix-factorization recommendation
Jaya Kawale, Hung Bui, Branislav Kveton, Long Tran-Thanh, and Sanjay Chawla · 2015
Cited alongside, same era.
Bootstrapped Thompson sampling and deep exploration
Ian Osband and Benjamin Van Roy · 2015
Cited alongside, same era.
Further optimal regret bounds for Thompson sampling
Shipra Agrawal and Navin Goyal
Cited in the paper.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal
Cited in the paper.
Perturbed-history exploration in stochastic multi-armed bandits
Branislav Kveton, Csaba Szepesvari, Mohammad Ghavamzadeh, and Craig Boutilier
Cited in the paper.
Later among the works it cites.
Deep Bayesian bandits showdown: An empirical comparison of Bayesian deep networks for Thompson sampling
Carlos Riquelme, George Tucker, and Jasper Snoek · 2018
Closest in time.
New insights into bootstrapping for bandits
Sharan Vaswani, Branislav Kveton, Zheng Wen, Anup Rao, Mark Schmidt, and Yasin Abbasi-Yadkori · 2018
Closest in time.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvari · 2019
Closest in time.