Fetching the paper…
Reading the bibliography…
Thompson sampling (TS) is a class of algorithms for sequential decision-making, which requires maintaining a posterior distribution over a model.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Riemannian geometry
Manfredo Perdigão do Carmo · 1992
Earlier work this paper cites.
Sample mean based index policies by o (log n) regret for the multi-armed bandit problem
Rajeev Agrawal · 1995
Earlier work this paper cites.
Error analysis for implicit approximations to solutions to cauchy problems
Jim Rulla · 1996
Earlier work this paper cites.
The variational formulation of the fokker–planck equation
Richard Jordan, David Kinderlehrer, and Felix Otto · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Gradient Flows in Metric Spaces and in the Space of Probability Measures
L. Ambrosio, N. Gigli, and G. Savaré · 2005
Earlier work this paper cites.
Optimal transport: old and new
C. Villani · 2008
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Cited alongside, same era.
An empirical evaluation of thompson sampling
Olivier Chapelle and Lihong Li · 2011
Cited alongside, same era.
Analysis of thompson sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal · 2012
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck, Nicolo Cesa-Bianchi, et al · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Weight uncertainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Cited alongside, same era.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Later among the works it cites.
Efficient exploration through bayesian deep q-networks
Kamyar Azizzadenesheli, Emma Brunskill, and Animashree Anandkumar · 2018
Later among the works it cites.
A unified particle-optimization framework for scalable bayesian sampling
Changyou Chen, Ruiyi Zhang, Wenlin Wang, Bai Li, and Liqun Chen · 2018
Later among the works it cites.
Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling
Carlos Riquelme, George Tucker, and Jasper Snoek · 2018
Later among the works it cites.
A tutorial on thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stein variational gradient descent: A general purpose bayesian inference algorithm
Qiang Liu and Dilin Wang · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
An information-theoretic analysis of thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Cited alongside, same era.
Boltzmann exploration done right
Nicolò Cesa-Bianchi, Claudio Gentile, Gábor Lugosi, and Gergely Neu · 2017
Cited alongside, same era.
First variation of the general curvature-dependent surface energy
Günay Do 𝒖 u
Cited in the paper.
Sharan Vaswani, Branislav Kveton, Zheng Wen, Anup Rao, Mark Schmidt, and Yasin Abbasi-Yadkori · 2018
Later among the works it cites.
Stochastic particle-optimization sampling and the non-asymptotic convergence theory
Jianyi Zhang, Ruiyi Zhang, and Changyou Chen · 2018
Later among the works it cites.
Policy optimization as wasserstein gradient flows
Ruiyi Zhang, Changyou Chen, Chunyuan Li, and Lawrence Carin · 2018
Later among the works it cites.
Learning structural weight uncertainty for sequential decision-making
Ruiyi Zhang, Chunyuan Li, Changyou Chen, and Lawrence Carin · 2018
Later among the works it cites.