Fetching the paper…
Reading the bibliography…
Thompson sampling is one of the most widely used algorithms for many online decision problems, due to its simplicity in implementation and superior empirical performance over other state-of-the-art methods.
Old dog learns new tricks: Randomized ucb for bandit problems
Vaswani, S · 1910
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Handbook of mathematical functions with formulas, graphs, and mathematical table
Abramowitz, M · 1965
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L · 1985
Earlier work this paper cites.
Sequential choice from several populations
Katehakis, M. N · 1995
Earlier work this paper cites.
Double explore-then-commit: Asymptotic optimality and beyond
Jin, T · 2002
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
Audibert, J.-Y · 2009
Earlier work this paper cites.
Ucb revisited: Improved regret bounds for the stochastic multi-armed bandit problem
Auer, P · 2010
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Chapelle, O · 2011
Earlier work this paper cites.
The kl-ucb algorithm for bounded stochastic bandits and beyond
Garivier, A · 2011
Earlier work this paper cites.
A finite-time analysis of multi-armed bandits problems with kullback-leibler divergences
Maillard, O.-A · 2011
Cited alongside, same era.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal, S · 2012
Cited alongside, same era.
Thompson sampling: An asymptotically optimal finite-time analysis
Kaufmann, E · 2012
Cited alongside, same era.
Open problem: Regret bounds for thompson sampling
Li, L · 2012
Cited alongside, same era.
Further optimal regret bounds for thompson sampling
Agrawal, S · 2013
Cited alongside, same era.
Prior-free and prior-dependent regret bounds for thompson sampling
Bubeck, S · 2013
Cited alongside, same era.
On bayesian index policies for sequential resource allocation
Kaufmann, E · 2016
Later among the works it cites.
Batched bandit problems
Perchet, V · 2016
Later among the works it cites.
Near-optimal regret bounds for thompson sampling
Agrawal, S · 2017
Later among the works it cites.
A minimax and asymptotically optimal algorithm for stochastic bandits
Ménard, P · 2017
Later among the works it cites.
Refining the confidence level for optimistic bandit strategies
Lattimore, T · 2018
Later among the works it cites.
Bandits with delayed, aggregated anonymous feedback
Pike-Burke, C · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thompson sampling for 1-dimensional exponential family bandits
Korda, N · 2013
Cited alongside, same era.
Learning to optimize via posterior sampling
Russo, D · 2014
Cited alongside, same era.
Optimally confident ucb: Improved regret for finite-armed bandits
Lattimore, T · 2015
Cited alongside, same era.
On explore-then-commit strategies
Garivier, A · 2016
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Auer, P
Cited in the paper.
The nonstochastic multiarmed bandit problem
Auer, P
Cited in the paper.
A tutorial on thompson sampling
Russo, D. J · 2018
Later among the works it cites.
Thompson sampling for combinatorial semi-bandits
Wang, S · 2018
Later among the works it cites.
Batched multi-armed bandits problem
Gao, Z · 2019
Later among the works it cites.
Bandit algorithms
Lattimore, T · 2020
Closest in time.