Fetching the paper…
Reading the bibliography…
We consider optimal sequential allocation in the context of the so-called stochastic multi-armed bandit model.
Audibert, Jean-YvesJ.-Y., Munos, RémiR. andSzepesvári, CsabaC. (2009). Exploration-exploitation tradeoff using variance estimates in multi-armed bandits. Theoret. Comput. Sci. 410 1876–1902
1902
Earlier work this paper cites.
Thompson, W. R.W. R. (1933). On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika 25 285–294
1933
Earlier work this paper cites.
Thompson, William R.W. R. (1935). On the Theory of Apportionment. Amer. J. Math. 57 450–456
1935
Earlier work this paper cites.
Wald, A.A. (1945). Sequential tests of statistical hypotheses. Ann. Math. Statist. 16 117–186
1945
Earlier work this paper cites.
Robbins, HerbertH. (1952). Some aspects of the sequential design of experiments. Bull. Amer. Math. Soc. (N.S.) 58 527–535
1952
Earlier work this paper cites.
Hoeffding, WassilyW. (1963). Probability inequalities for sums of bounded random variables. J. Amer. Statist. Assoc. 58 13–30
1963
Earlier work this paper cites.
Gittins, J. C.J. C. (1979). Bandit processes and dynamic allocation indices (with discussion). J. R. Stat. Soc. Ser. B Stat. Methodol. 41 148–177
1979
Earlier work this paper cites.
1980
Earlier work this paper cites.
Lai, T. L.T. L. andRobbins, HerbertH. (1985). Asymptotically efficient adaptive allocation rules. Adv. in Appl. Math. 6 4–22
1985
Earlier work this paper cites.
Chang, FuF. andLai, Tze LeungT. L. (1987). Optimal stopping and dynamic allocation. Adv. in Appl. Probab. 19 829–853
1987
Earlier work this paper cites.
Chow, Yuan ShihY. S. andTeicher, HenryH. (1988). Probability Theory: Independence, Interchangeability, Martingales, 2nd ed. Springer, New York
1988
Earlier work this paper cites.
Weber, RichardR. (1992). On the Gittins index for multiarmed bandits. Ann. Appl. Probab. 2 1024–1033
1992
Earlier work this paper cites.
Agrawal, RajeevR. (1995). Sample mean based index policies with O ( log n ) O(\log n) regret for the multi-armed bandit problem. Adv. in Appl. Probab. 27 1054–1078
1995
Cited alongside, same era.
Burnetas, Apostolos N.A. N. andKatehakis, Michael N.M. N. (1996). Optimal adaptive policies for sequential allocation problems. Adv. in Appl. Math. 17 122–142
1996
Cited alongside, same era.
Burnetas, Apostolos N.A. N. andKatehakis, Michael N.M. N. (1997). Optimal adaptive policies for Markov decision processes. Math. Oper. Res. 22 222–255
1997
Cited alongside, same era.
Dembo, AmirA. andZeitouni, OferO. (1998). Large Deviations Techniques and Applications, 2nd ed. Applications of Mathematics (New York) 38. Springer, New York
1998
Cited alongside, same era.
Lehmann, E. L.E. L. andCasella, GeorgeG. (1998). Theory of Point Estimation, 2nd ed. Springer, New York
Filippi, S.S., Cappé, O.O. andGarivier, A.A. (2010). Optimism in reinforcement learning and Kullback–Leibler divergence. In Proceedings of the 48th Annual Allerton Conference on Communication, Control, and Computing. IEEE Press, Piscataway, NJ
2010
Later among the works it cites.
Honda, J.J. andTakemura, A.A. (2010). An asymptotically optimal bandit algorithm for bounded support models. In Proceedings of the 23rd Annual Conference on Learning Theory. Omnipress, Madison, WI
2010
Later among the works it cites.
Garivier, A.A. andCappé, O.O. (2011). The KL-UCB algorithm for bounded stochastic bandits and beyond. In Proceedings of the 24th Annual Conference on Learning Theory. JMLR C&WP
2011
Later among the works it cites.
Gittins, J.J., Glazebrook, K.K. andWeber, R.R. (2011). Multi-Armed Bandit Allocation Indices. Wiley, New York
2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
1998
Cited alongside, same era.
van der Vaart, A. W.A. W. (2000). Asymptotic Statistics. Cambridge Univ. Press, Cambridge
2000
Cited alongside, same era.
Owen, A. B.A. B. (2001). Empirical Likelihood. Chapman & Hall/CRC, Boca Raton, FL
2001
Cited alongside, same era.
Auer, P.P., Cesa-Bianchi, N.N. andFischer, P.P. (2002). Finite-time analysis of the multiarmed bandit problem. Machine Learning 47 235–256
2002
Cited alongside, same era.
Burnetas, Apostolos N.A. N. andKatehakis, Michael N.M. N. (2003). Asymptotic Bayes analysis for the finite-horizon one-armed-bandit problem. Probab. Engrg. Inform. Sci. 17 53–82
2003
Cited alongside, same era.
Massart, PascalP. (2007). Concentration Inequalities and Model Selection. Lecture Notes in Math. 1896. Springer, Berlin
2007
Cited alongside, same era.
Wainwright, Martin J.M. J. andJordan, Michael I.M. I. (2008). Graphical models, exponential families, and variational inference. Foundation and Trends in Machine Learning 1 1–305
2008
Cited alongside, same era.
Audibert, Jean-YvesJ.-Y. andBubeck, SébastienS. (2010). Regret bounds and minimax policies under partial monitoring. J. Mach. Learn. Res. 11 2785–2836
2010
Cited alongside, same era.
Honda, J.J. andTakemura, A.A. (2011). An asymptotically optimal policy for finite support models in the multiarmed bandit problem. Machine Learning 85 361–391
2011
Later among the works it cites.
Maillard, O-A.O.-A., Munos, R.R. andStoltz, G.G. (2011). A finite-time analysis of multi-armed bandits problems with Kullback–Leibler divergences. In Proceedings of the 24th Annual Conference on Learning Theory. JMLR C&WP
2011
Later among the works it cites.
Bubeck, S.S. andCesa-Bianchi, N.N. (2012). Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Foundations and Trends in Machine Learning 5 1–122
2012
Closest in time.
Cappé, OlivierO., Garivier, AurelienA. andKaufmann, EmilieE. (2012). py/maBandits: Matlab and Python packages for multi-armed bandits. Available at http://mloss.org/software/ view/415/
2012
Closest in time.
2012
Closest in time.
Kaufmann, E.E., Cappé, O.O. andGarivier, A.A. (2012). On Bayesian upper confidence bounds for bandit problems. In Proceedings of the 15th International Conference on Artificial Intelligence and Statistics 22 592–600. JMLR W&CP
2012
Closest in time.
Kaufmann, E.E., Korda, N.N. andMunos, R.R. (2012). Thompson sampling: An asymptotically optimal finite time analysis. In Proceedings of the 23rd International Conference on Algorithmic Learning Theory 199–213. Springer, New York
2012
Closest in time.
Cappé, O.O., Garivier, A.A., Maillard, O-A.O.-A., Munos, R.R. andStoltz, G.G. (2013). Supplement to “Kullback–Leibler upper confidence bounds for optimal sequential allocation.” DOI: \doiurl
2013
Closest in time.