Fetching the paper…
Reading the bibliography…
Thompson Sampling is one of the oldest heuristics for multi-armed bandit problems.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, William R · 1933
Earlier work this paper cites.
Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables
Abramowitz, Milton and Stegun, Irene A · 1964
Earlier work this paper cites.
A one-armed bandit problem with a concomitant variable
Woodroofe, Michael · 1979
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L. and Robbins, H · 1985
Earlier work this paper cites.
One-armed badit problem with covariates
Sarkar, Jyotirmoy · 1991
Earlier work this paper cites.
Associative Reinforcement Learning: Functions in k-DNF
Kaelbling, Leslie Pack · 1994
Earlier work this paper cites.
Exploration and Inference in Learning from Reinforcement
Wyatt, Jeremy · 1997
Earlier work this paper cites.
A Bayesian Framework for Reinforcement Learning
Strens, Malcolm J. A · 2000
Earlier work this paper cites.
Using Confidence Bounds for Exploitation-Exploration Trade-offs
Auer, Peter · 2002
Earlier work this paper cites.
The Nonstochastic Multiarmed Bandit Problem
Auer, Peter, Cesa-Bianchi, Nicolò, Freund, Yoav, and Schapire, Robert E · 2002
Cited alongside, same era.
Experience-efficient learning in associative bandit problems
Strehl, Alexander L., Mesterharm, Chris, Littman, Michael L., and Hirsh, Haym · 2006
Cited alongside, same era.
The Epoch-Greedy Algorithm for Multi-armed Bandits with Side Information
Langford, John and Zhang, Tong · 2007
Cited alongside, same era.
Stochastic Linear Optimization under Bandit Feedback
Dani, Varsha, Hayes, Thomas P., and Kakade, Sham M · 2008
Cited alongside, same era.
Parametric Bandits: The Generalized Linear Case
Filippi, Sarah, Cappé, Olivier, Garivier, Aurélien, and Szepesvári, Csaba · 2010
Cited alongside, same era.
Web-Scale Bayesian Click-Through rate Prediction for Sponsored Search Advertising in Microsoft’s Bing Search Engine
Graepel, Thore, Candela, Joaquin Quiñonero, Borchert, Thomas, and Herbrich, Ralf · 2010
Improved Algorithms for Linear Stochastic Bandits
Abbasi-Yadkori, Yasin, Pál, Dávid, and Szepesvári, Csaba · 2011
Later among the works it cites.
An Empirical Evaluation of Thompson Sampling
Chapelle, Olivier and Li, Lihong · 2011
Later among the works it cites.
Contextual Bandits with Linear Payoff Functions
Chu, Wei, Li, Lihong, Reyzin, Lev, and Schapire, Robert E · 2011
Later among the works it cites.
Simulation studies in optimistic Bayesian sampling in contextual-bandit problems
May, Benedict C. and Leslie, David S · 2011
Later among the works it cites.
Optimistic Bayesian sampling in contextual-bandit problems
May, Benedict C., Korda, Nathan, Lee, Anthony, and Leslie, David S · 2011
Later among the works it cites.
Analysis of Thompson Sampling for the Multi-armed Bandit Problem
Agrawal, Shipra and Goyal, Navin · 2012
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Solving Two-Armed Bernoulli Bandit Problems Using a Bayesian Learning Automaton
Granmo, O.-C · 2010
Cited alongside, same era.
Linearly Parametrized Bandits
Ortega, Pedro A. and Braun, Daniel A · 2010
Cited alongside, same era.
A modern Bayesian look at the multi-armed bandit
Scott, S · 2010
Cited alongside, same era.
Thompson Sampling for Contextual Bandits with Linear Payoffs
Agrawal, Shipra and Goyal, Navin
Cited in the paper.
Further Optimal Regret Bounds for Thompson Sampling
Agrawal, Shipra and Goyal, Navin
Cited in the paper.
Towards minimax policies for online linear optimization with bandit feedback
Bubeck, Sébastien, Cesa-Bianchi, Nicolò, and Kakade, Sham M · 2012
Closest in time.
Open Problem: Regret Bounds for Thompson Sampling
Chapelle, Olivier and Li, Lihong · 2012
Closest in time.
Thompson Sampling: An Optimal Finite Time Analysis
Kaufmann, Emilie, Korda, Nathaniel, and Munos, Rémi · 2012
Closest in time.