Fetching the paper…
Reading the bibliography…
This paper considers the use of a simple posterior sampling algorithm to balance between exploration and exploitation when learning to optimize actions such as in multi-armed bandit problems.
Computationally related problems
Sahni, A. 1974 · 1974
Earlier work this paper cites.
A dynamic allocation index for the discounted multiarmed bandit problem
Gittins, J.C., D.M. Jones. 1979 · 1979
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T.L., H. Robbins. 1985 · 1985
Earlier work this paper cites.
Adaptive treatment allocation and the multi-armed bandit problem
Lai, T.L. 1987 · 1987
Earlier work this paper cites.
Theory of point estimation, vol. 31
Lehmann, E.L., G. Casella. 1998 · 1998
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., N. Cesa-Bianchi, P. Fischer. 2002 · 2002
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V., T.P. Hayes, S.M. Kakade. 2008 · 2008
Earlier work this paper cites.
Multi-armed bandits in metric spaces
Kleinberg, R., A. Slivkins, E. Upfal. 2008 · 2008
Earlier work this paper cites.
Forced-exploration based algorithms for playing in stochastic linear bandits
Abbasi-Yadkori, Y., A. Antos, C. Szepesvári. 2009 · 2009
Earlier work this paper cites.
Minimax policies for bandits games
Audibert, J.-Y., S. Bubeck. 2009 · 2009
Earlier work this paper cites.
X-armed bandits
Bubeck, S., R. Munos, G. Stoltz, Cs. Szepesvári. 2011 · 2010
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Filippi, S., O. Cappé, A. Garivier, C. Szepesvári. 2010 · 2010
Cited alongside, same era.
Linearly parameterized bandits
Rusmevichientong, P., J.N. Tsitsiklis. 2010 · 2010
Cited alongside, same era.
A modern Bayesian look at the multi-armed bandit
Scott, S.L. 2010 · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., D. Pál, C. Szepesvári. 2011 · 2011
Cited alongside, same era.
Bandits, query learning, and the haystack dimension
Amin, K., M. Kearns, U. Syed. 2011 · 2011
Cited alongside, same era.
Contextual bandit algorithms with supervised learning guarantees
Beygelzimer, A., J. Langford, L. Li, L. Reyzin, R.E. Schapire. 2011 · 2011
Cited alongside, same era.
Linear bandits in high dimension and recommendation systems
D., Yash, A. Montanari. 2012 · 2012
Later among the works it cites.
Thompson sampling: an asymptotically optimal finite time analysis
Kauffmann, E., N. Korda, R. Munos. 2012 · 2012
Later among the works it cites.
Open problem: Regret bounds for Thompson sampling
Li, L., O. Chapelle. 2012 · 2012
Later among the works it cites.
The knowledge gradient algorithm for a general class of online learning problems
Ryzhov, I.O., W.B. Powell, P.I. Frazier. 2012 · 2012
Later among the works it cites.
Prior-free and prior-dependent regret bounds for thompson sampling
Bubeck, S., C.-Y. Liu. 2013 · 2013
Closest in time.
Kullback-Leibler upper confidence bounds for optimal sequential allocation
Cappé, O., A. Garivier, O.-A. Maillard, R. Munos, G. Stoltz. 2013 · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An empirical evaluation of Thompson sampling
Chapelle, O., L. Li. 2011 · 2011
Cited alongside, same era.
Multi-Armed Bandit Allocation Indices
Gittins, J., K. Glazebrook, R. Weber. 2011 · 2011
Cited alongside, same era.
Information-theoretic regret bounds for Gaussian process optimization in the bandit setting
Srinivas, N., A. Krause, S.M. Kakade, M. Seeger. 2012 · 2011
Cited alongside, same era.
Online-to-confidence-set conversions and application to sparse stochastic bandits
Abbasi-Yadkori, Y., D. Pal, C. Szepesvári. 2012 · 2012
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S., N. Cesa-Bianchi. 2012 · 2012
Cited alongside, same era.
Analysis of Thompson sampling for the multi-armed bandit problem
Agrawal, S., N. Goyal. 2012a
Cited in the paper.
Closest in time.
Thompson sampling for complex bandit problems
Gopalan, A., S. Mannor, Y. Mansour. 2013 · 2013
Closest in time.
Thompson sampling for one-dimensional exponential family bandits
Korda, N., E. Kaufmann, R. Munos. 2013 · 2013
Closest in time.
Generalized thompson sampling for contextual bandits
Li, Lihong. 2013 · 2013
Closest in time.
Stochastic simultaneous optimistic optimization
Valko, M., A. Carpentier, R. Munos. 2013 · 2013
Closest in time.
Optimistic Bayesian sampling in contextual-bandit problems
May, B.C., N. Korda, A. Lee, D.S. Leslie. 2012 · 2069
Closest in time.