Cambridge university press, 2006
N. Cesa-Bianchi and G. Lugosi, Prediction, learning, and games · 2006
Cited alongside, same era.
V. Dani, T. P. Hayes, and S. M. Kakade, “Stochastic linear optimization under bandit feedback.,” in COLT
2008
Cited alongside, same era.
S. L. Scott, “A modern bayesian look at the multi-armed bandit,” Applied Stochastic Models in Business and Industry
2010
Cited alongside, same era.
Y. Abbasi-Yadkori and C. Szepesvári, “Regret bounds for the adaptive control of linear quadratic systems,” in Proceedings of the 24th Annual Conference on Learning Theory
2011
Cited alongside, same era.
O. Chapelle and L. Li, “An empirical evaluation of thompson sampling,” in NIPS
2011
Cited alongside, same era.
M. Ibrahimi, A. Javanmard, and B. V. Roy, “Efficient reinforcement learning for high dimensional linear quadratic systems,” in Advances in Neural Information Processing Systems (NIPS)
2012
Cited alongside, same era.
S. Agrawal and N. Goyal, “Analysis of thompson sampling for the multi-armed bandit problem,” in Conference on Learning Theory
2012
Cited alongside, same era.
E. Kaufmann, N. Korda, and R. Munos, “Thompson sampling: An asymptotically optimal finite-time analysis,” in International Conference on Algorithmic Learning Theory
2012
Cited alongside, same era.
S. Agrawal and N. Goyal, “Thompson sampling for contextual bandits with linear payoffs.,” in ICML (3)
2013
Cited alongside, same era.
I. Osband, D. Russo, and B. Van Roy, “(More) efficient reinforcement learning via posterior sampling,” in NIPS
2013
Cited alongside, same era.
available at: www.cs.cmu.edu/ gautamd/Files/maxGaussians.pdf
Original
G. Dasarathy, “A simple probability trick for bounding the expected maximum of n random variables,”
Cited in the paper.