Fetching the paper…
Reading the bibliography…
Thompson Sampling has been widely used for contextual bandit problems due to the flexibility of its modeling power.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Associative reinforcement learning using linear probabilistic concepts
Naoki Abe and Philip M Long · 1999
Earlier work this paper cites.
Averaging expert predictions
Jyrki Kivinen and Manfred K Warmuth · 1999
Earlier work this paper cites.
Information-theoretic determination of minimax rates of convergence
Yuhong Yang and Andrew Barron · 1999
Earlier work this paper cites.
Competitive on-line statistics
Volodya Vovk · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
Olivier Chapelle and Lihong Li · 2011
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient Langevin dynamics
Max Welling and Yee W Teh · 2011
Cited alongside, same era.
Contextual bandit learning with predictable rewards
Alekh Agarwal, Miroslav Dudík, Satyen Kale, John Langford, and Robert Schapire · 2012
Cited alongside, same era.
Further optimal regret bounds for Thompson sampling
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Bounded regret for finite-armed structured bandits
Tor Lattimore and Rémi Munos · 2014
Cited alongside, same era.
From ads to interventions: Contextual bandits in mobile health
Ambuj Tewari and Susan A Murphy · 2017
Later among the works it cites.
Improved regret bounds for thompson sampling in linear quadratic control problems
Marc Abeille and Alessandro Lazaric · 2018
Later among the works it cites.
On oracle-efficient PAC RL with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2018
Later among the works it cites.
Deep exploration via randomized value functions
Ian Osband, Benjamin Van Roy, Daniel J Russo, and Zheng Wen · 2019
Later among the works it cites.
Improved worst-case regret bounds for randomized least-squares value iteration
Priyank Agrawal, Jinglin Chen, and Nan Jiang · 2020
Later among the works it cites.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Latent bandits
Odalric-Ambrym Maillard and Shie Mannor · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Making contextual decisions with low technical debt
Alekh Agarwal, Sarah Bird, Markus Cozowicz, Luong Hoang, John Langford, Stephen Lee, Jiaji Li, Dan Melamed, Gal Oshri, Oswaldo Ribas, et al · 2016
Cited alongside, same era.
An information-theoretic analysis of Thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Dylan Foster and Alexander Rakhlin · 2020
Later among the works it cites.
Adapting to misspecification in contextual bandits
Dylan J Foster, Claudio Gentile, Mehryar Mohri, and Julian Zimmert · 2020
Later among the works it cites.
Joey Hong, Branislav Kveton, Manzil Zaheer, Yinlam Chow, Amr Ahmed, and Craig Boutilier · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
David Simchi-Levi and Yunzong Xu · 2020
Later among the works it cites.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 2020
Later among the works it cites.
Mirror descent and the information ratio
Tor Lattimore and Andras Gyorgy · 2021
Closest in time.