Fetching the paper…
Reading the bibliography…
Thompson sampling is an algorithm for online decision problems where actions are taken sequentially in a manner that must balance between exploiting what is known to maximize immediate performance and investing to accumulate new information that may improve future performance.
“On the likelihood that one unknown probability exceeds another in view of the evidence of two samples”
W.R. Thompson · 1933
Earlier work this paper cites.
“On the theory of apportionment”
William Thompson · 1935
Earlier work this paper cites.
“A dynamic allocation index for the discounted multiarmed bandit problem”
J.C. Gittins and D.M. Jones · 1979
Earlier work this paper cites.
“Asymptotically efficient adaptive allocation rules”
T.L. Lai and H. Robbins · 1985
Earlier work this paper cites.
“The multi-armed bandit problem: decomposition and computation”
Michael. Katehakis and Arthur. Veinott Jr · 1987
Earlier work this paper cites.
“Explaining the Gibbs sampler”
George Casella and Edward George · 1992
Earlier work this paper cites.
“Exponential convergence of Langevin distributions and their discrete approximations”
Gareth Roberts and Richard Tweedie · 1996
Earlier work this paper cites.
“Exploration and inference in learning from reinforcement”, 1997
Jeremy Wyatt · 1997
Earlier work this paper cites.
“Optimal scaling of discrete approximations to Langevin diffusions”
Gareth Roberts and Jeffrey Rosenthal · 1998
Earlier work this paper cites.
“Reinforcement learning: An introduction”
Richard Sutton and Andrew Barto · 1998
Earlier work this paper cites.
“A Bayesian framework for reinforcement learning”
Malcolm Strens · 2000
Earlier work this paper cites.
“Finite-time analysis of the multiarmed bandit problem”
Peter Auer, Nicolò Cesa-Bianchi and Paul Fischer · 2002
Earlier work this paper cites.
“Ergodicity for SDEs and approximations: locally Lipschitz vector fields and degenerate noise”
Jonathan Mattingly, Andrew Stuart and Desmond Higham · 2002
Earlier work this paper cites.
“An experimental comparison of click position-bias models”
Nick Craswell, Onno Zoeter, Michael Taylor and Bill Ramsey · 2008
Earlier work this paper cites.
“Stochastic linear optimization under bandit feedback”
V. Dani, T.P. Hayes and S.M. Kakade · 2008
Earlier work this paper cites.
“A knowledge-gradient policy for sequential information collection”
P.I. Frazier, W.B. Powell and S. Dayanik · 2008
Earlier work this paper cites.
“Multi-armed bandits in metric spaces”
R. Kleinberg, A. Slivkins and E. Upfal · 2008
Earlier work this paper cites.
“The knowledge-gradient policy for correlated normal beliefs”
P. Frazier, W. Powell and S. Dayanik · 2009
Earlier work this paper cites.
“Web-scale Bayesian click-through rate prediction for sponsored search advertising in Microsoft’s Bing search engine”
T. Graepel, J.Q. Candela, T. Borchert and R. Herbrich · 2010
Earlier work this paper cites.
“Near-optimal regret bounds for reinforcement learning”
T. Jaksch, R. Ortner and P. Auer · 2010
Earlier work this paper cites.
“A Contextual-bandit approach to personalized news article recommendation”
Lihong Li, Wei Chu, John Langford and Robert. Schapire · 2010
Earlier work this paper cites.
“Linearly parameterized bandits”
P. Rusmevichientong and J.N. Tsitsiklis · 2010
Earlier work this paper cites.
“A modern Bayesian look at the multi-armed bandit”
S.L. Scott · 2010
Earlier work this paper cites.
“Improved algorithms for linear stochastic bandits”
Yasin Abbasi-Yadkori, Dávid Pál and Csaba Szepesvári · 2011
Earlier work this paper cites.
“X-armed bandits”
S. Bubeck, R. Munos, G. Stoltz and C. Szepesvári · 2011
Earlier work this paper cites.
“An empirical evaluation of Thompson sampling”
O. Chapelle and L. Li · 2011
Earlier work this paper cites.
“Multi-armed bandit allocation indices”
John Gittins, Kevin Glazebrook and Richard Weber · 2011
Earlier work this paper cites.
“Bayesian learning via stochastic gradient Langevin dynamics”
Max Welling and Yee Teh · 2011
Earlier work this paper cites.
“Analysis of Thompson sampling for the multi-armed bandit problem”
Shipra Agrawal and Navin Goyal · 2012
Cited alongside, same era.
“Regret analysis of stochastic and nonstochastic multi-armed bandit problems”
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Cited alongside, same era.
“Thompson sampling: an asymptotically optimal finite time analysis”
E. Kauffmann, N. Korda and R. Munos · 2012
Cited alongside, same era.
“On Bayesian upper confidence bounds for bandit problems”
E. Kaufmann, O. Cappé and A. Garivier · 2012
Cited alongside, same era.
“Information-Theoretic regret bounds for Gaussian process optimization in the bandit setting”
N. Srinivas, A. Krause, S.M. Kakade and M. Seeger · 2012
Cited alongside, same era.
“Computational advertising: the LinkedIn way”
Deepak Agarwal · 2013
Cited alongside, same era.
“Multi-scale exploration of convex functions and bandit convex optimization”
Sébastien Bubeck and Ronen Eldan · 2016
Later among the works it cites.
“Sampling from strongly log-concave distributions with the Unadjusted Langevin Algorithm”
Alain Durmus and Eric Moulines · 2016
Later among the works it cites.
“Online algorithms for parameter mean and variance estimation in dynamic regression”
Carlos. Gómez-Uribe · 2016
Later among the works it cites.
“Deep exploration via bootstrapped DQN”
Ian Osband, Charles Blundell, Alexander Pritzel and Benjamin Van · 2016
Later among the works it cites.
“Generalization and exploration via randomized value functions”
Ian Osband, Benjamin Van and Zheng Wen · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Further optimal regret bounds for Thompson sampling”
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
“Thompson sampling for contextual bandits with linear payoffs”
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
“Bayesian mixture modelling and inference based Thompson sampling in Monte-Carlo tree search”
Aijun Bai, Feng Wu and Xiaoping Chen · 2013
Cited alongside, same era.
“Kullback-Leibler upper confidence bounds for optimal sequential allocation”
O. Cappé et al · 2013
Cited alongside, same era.
“(More) Efficient reinforcement learning via posterior sampling”
I. Osband, D. Russo and B. Van · 2013
Cited alongside, same era.
“Eluder Dimension and the Sample Complexity of Optimistic Exploration”
D. Russo and B. Van · 2013
Cited alongside, same era.
“An Information-Theoretic analysis of Thompson sampling”
D. Russo and B. Van · 2016
Later among the works it cites.
“Simple bayesian algorithms for best arm identification”
Daniel Russo · 2016
Later among the works it cites.
“Consistency and fluctuations for stochastic gradient Langevin dynamics”
Yee Teh, Alexandre Thiery and Sebastian Vollmer · 2016
Later among the works it cites.
“Linear Thompson sampling revisited”
Marc Abeille and Alessandro Lazaric · 2017
Closest in time.
“Thompson sampling for the MNL-bandit”
Shipra Agrawal, Vashist Avadhanula, Vineet Goyal and Assaf Zeevi · 2017
Closest in time.
“Choosing a Good Toolkit: Bayes-Rule Based Heuristics”
Alejandro Francetich and David. Kreps · 2017
Closest in time.
“Choosing a Good Toolkit: Reinforcement Learning”
Alejandro Francetich and David. Kreps · 2017
Closest in time.
“An efficient bandit algorithm for realtime multivariate optimization”
Daniel. Hill et al · 2017
Closest in time.
“Thompson sampling for stochastic control: the finite parameter case”
Michael Kim · 2017
Closest in time.
“Information directed sampling for stochastic bandits with graph feedback”
Fang Liu, Swapna Buccapatnam and Ness Shroff · 2017
Closest in time.
“Ensemble Sampling”
Xiuyuan Lu and Benjamin Van · 2017
Closest in time.
“Deep exploration via randomized value functions”
Ian Osband, Daniel Russo, Zheng Wen and Benjamin Van · 2017
Closest in time.
“On optimistic versus randomized exploration in reinforcement learning”
Ian Osband and Benjamin Van · 2017
Closest in time.
“Why is posterior sampling better than optimism for reinforcement learning?”
Ian Osband and Benjamin Van · 2017
Closest in time.
“Learning unknown Markov decision processes: A Thompson sampling approach”
Yi Ouyang, Mukul Gagrani, Ashutosh Nayyar and Rahul Jain · 2017
Closest in time.
“Customer acquisition via display advertising using multi-armed bandit experiments”
Eric Schwartz, Eric Bradlow and Peter Fader · 2017
Closest in time.
“Exploiting the natural exploration in contextual bandits”
Hamsa Bastani, Mohsen Bayati and Khashayar Khosravi · 2018
Closest in time.
“Sampling from a log-concave distribution with projected Langevin Monte Carlo”
Sébastien Bubeck, Ronen Eldan and Joseph Lehec · 2018
Closest in time.
“Convergence of Langevin MCMC in KL-divergence”
Xiang Cheng and Peter Bartlett · 2018
Closest in time.
“Coordinated exploration in concurrent reinforcement learning”
Maria Dimakopoulou and Benjamin Van · 2018
Closest in time.
“Parallelised Bayesian optimisation via Thompson sampling”
Kirthevasan Kandasamy, Akshay Krishnamurthy, Jeff Schneider and Barnabas Poczos · 2018
Closest in time.
“Learning to optimize via information-directed sampling”
Daniel Russo and Benjamin Van · 2018
Closest in time.
“Satisficing in time-sensitive bandit learning”
Daniel Russo and Benjamin Van · 2018
Closest in time.