Fetching the paper…
Reading the bibliography…
Thompson sampling provides a solution to bandit problems in which new observations are allocated to arms with the posterior probability that an arm is optimal.
Exploration-exploitation tradeoff using variance estimates in multi-armed bandits
Audibert, J.-Y., Munos, R., and Szepesvári, C. (2009) · 1902
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. (1933) · 1933
Earlier work this paper cites.
Bayesian inference and the parametric bootstrap
Efron, B. (2012) · 1971
Earlier work this paper cites.
Bootstrap methods: Another look at the jackknife
Efron, B. (1979) · 1979
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
Gittins, J. C. (1979) · 1979
Earlier work this paper cites.
Multi-armed bandits and the Gittins index
Whittle, P. (1980) · 1980
Earlier work this paper cites.
Bootstrapping regression models
Freedman, D. A. (1981) · 1981
Earlier work this paper cites.
The Bayesian bootstrap
Rubin, D. B. (1981) · 1981
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L. and Robbins, H. (1985) · 1985
Cited alongside, same era.
Comment on “Approximate Bayesian inference with the weighted likelihood bootstrap”
Liu, J. S. and Rubin, D. B. (1994) · 1994
Cited alongside, same era.
Approximate Bayesian inference with the weighted likelihood bootstrap
Newton, M. A. and Raftery, A. E. (1994) · 1994
Cited alongside, same era.
Bandit problems and the exploration/exploitation tradeoff
Macready, W. G. and Wolpert, D. H. (1998) · 1998
Cited alongside, same era.
Lossless online Bayesian bagging
Lee, H. K. H. and Clyde, M. A. (2004) · 2004
Cited alongside, same era.
Online bagging and boosting
Oza, N. (2001) · 2005
Cited alongside, same era.
UCB revisited: Improved regret bounds for the stochastic multi-armed bandit problem
Auer, P. and Ortner, R. (2010) · 2010
Later among the works it cites.
Web-scale Bayesian click-through rate prediction for sponsored search advertising in Microsoft’s Bing search engine
Graepel, T., Candela, J. Q., Borchert, T., and Herbrich, R. (2010) · 2010
Later among the works it cites.
A modern Bayesian look at the multi-armed bandit
Scott, S. L. (2010) · 2010
Later among the works it cites.
An empirical evaluation of Thompson sampling
Chapelle, O. and Li, L. (2011) · 2011
Later among the works it cites.
The KL-UCB algorithm for bounded stochastic bandits and beyond
Garivier, A. and Cappé, O. (2011) · 2011
Later among the works it cites.
Thompson sampling: An asymptotically optimal finite-time analysis
Kaufmann, E., Korda, N., and Munos, R. (2012) · 2012
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Map-Reduce for Machine Learning on Multicore
Chu, C.-t., Kim, S. K., Lin, Y.-a., and Ng, A. Y. (2007) · 2007
Cited alongside, same era.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
Hastie, T., Tibshirani, R., and Friedman, J. (2008) · 2008
Cited alongside, same era.
Later among the works it cites.
Bootstrapping data arrays of arbitrary order
Owen, A. B. and Eckles, D. (2012) · 2012
Later among the works it cites.
Model-robust regression and a Bayesian “sandwich” estimator
Szpiro, A. A., Rice, K. M., and Lumley, T. (2010) · 2099
Closest in time.