Fetching the paper…
Reading the bibliography…
We consider a multi-armed bandit problem in a setting where each arm produces a noisy reward realization which depends on an observable random covariate.
Robbins, HerbertH. (1952). Some aspects of the sequential design of experiments. Bull. Amer. Math. Soc. (N.S.) 58 527–535
1952
Earlier work this paper cites.
Vogel, WalterW. (1960). An asymptotic minimax theorem for the two armed bandit problem. Ann. Math. Statist. 31 444–451
1960
Earlier work this paper cites.
Woodroofe, MichaelM. (1979). A one-armed bandit problem with a concomitant variable. J. Amer. Statist. Assoc. 74 799–806
1979
Earlier work this paper cites.
Bather, J. A.J. A. (1981). Randomized allocation of treatments in sequential experiments. J. R. Stat. Soc. Ser. B Stat. Methodol. 43 265–292
1981
Earlier work this paper cites.
Lai, T. L.T. L. andRobbins, HerbertH. (1985). Asymptotically efficient adaptive allocation rules. Adv. in Appl. Math. 6 4–22
1985
Earlier work this paper cites.
Mammen, EnnoE. andTsybakov, Alexandre B.A. B. (1999). Smooth discrimination analysis. Ann. Statist. 27 1808–1829
1999
Earlier work this paper cites.
Auer, P.P., Cesa-Bianchi, N.N. andFischer, P.P. (2002). Finite-time analysis of the multiarmed bandit problem. Mach. Learn. 47 235–256
2002
Earlier work this paper cites.
Yang, YuhongY. andZhu, DanD. (2002). Randomized allocation with nonparametric estimation for a multi-armed bandit problem with covariates. Ann. Statist. 30 100–121
2002
Earlier work this paper cites.
Tsybakov, Alexandre B.A. B. (2004). Optimal aggregation of classifiers in statistical learning. Ann. Statist. 32 135–166
2004
Earlier work this paper cites.
Audibert, J. Y.J. Y. andTsybakov, A. B. B.A. B. B. (2005). Fast learning rates for plug-in classifiers under the margin condition. Preprint, Laboratoire de Probabilités et Modèles Aléatoires, Univ. Paris VI and VII. Available at arXiv: \arxivurl
2005
Cited alongside, same era.
Wang, Chih-ChunC.-C., Kulkarni, Sanjeev R.S. R. andPoor, H. VincentH. V. (2005). Bandit problems with side observations. IEEE Trans. Automat. Control 50 338–355
2005
Cited alongside, same era.
Cesa-Bianchi, NicolòN. andLugosi, GáborG. (2006). Prediction, Learning, and Games. Cambridge Univ. Press, Cambridge
2006
Cited alongside, same era.
Even-Dar, EyalE., Mannor, ShieS. andMansour, YishayY. (2006). Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems. J. Mach. Learn. Res. 7 1079–1105
2006
Cited alongside, same era.
Langford, J.J. andZhang, T.T. (2008). The epoch-greedy algorithm for multi-armed bandits with side information. In Advances in Neural Information Processing Systems 20 (J. C.J. C. Platt, D.D. Koller, Y.Y. Singer andS.S. Roweis, eds.) 817–824. MIT Press, Cambridge, MA
2008
Later among the works it cites.
Goldenshluger, AlexanderA. andZeevi, AssafA. (2009). Woodroofe’s one-armed bandit problem revisited. Ann. Appl. Probab. 19 1603–1633
2009
Later among the works it cites.
Audibert, Jean-YvesJ.-Y. andBubeck, SébastienS. (2010). Regret bounds and minimax policies under partial monitoring. J. Mach. Learn. Res. 11 2785–2836
2010
Later among the works it cites.
Auer, PeterP. andOrtner, RonaldR. (2010). UCB revisited: Improved regret bounds for the stochastic multi-armed bandit problem. Period. Math. Hungar. 61 55–65
2010
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Audibert, Jean-YvesJ.-Y. andTsybakov, Alexandre B.A. B. (2007). Fast learning rates for plug-in classifiers. Ann. Statist. 35 608–633
2007
Cited alongside, same era.
Hazan, EladE. andMegiddo, NimrodN. (2007). Online learning with prior knowledge. In Learning Theory. Lecture Notes in Computer Science 4539 499–513. Springer, Berlin
2007
Cited alongside, same era.
Juditsky, A.A., Nazin, A. V.A. V., Tsybakov, A. B.A. B. andVayatis, N.N. (2008). Gap-free bounds for stochastic multi-armed bandit. In Proceedings of the 17th IFAC World Congress
2008
Cited alongside, same era.
Kakade, S.S., Shalev-Shwartz, S.S. andTewari, A.A. (2008). Efficient bandit algorithms for online multiclass prediction. In Proceedings of the 25th Annual International Conference on Machine Learning (ICML 2008) (AndrewA. McCallum andSamS. Roweis, eds.) 440–447. Omnipress, Helsinki, Finland
2008
Cited alongside, same era.
Lu, T.T., Pál, D.D. andPál, M.M. (2010). Showing relevant ads via Lipschitz context multi-armed bandits. JMLR: Workshop and Conference Proceedings 9 485–492
2010
Later among the works it cites.
Rigollet, P.P. andZeevi, A.A. (2010). Nonparametric bandits with covariates. In COLT (AdamA. Tauman Kalai andMehryarM. Mohri, eds.) 54–66. Omnipress, Haifa, Israel
2010
Later among the works it cites.
Goldenshluger, A.A. andZeevi, A.A. (2011). A note on performance limitations in bandit problems with side information. IEEE Trans. Inform. Theory 57 1707–1713
2011
Closest in time.
Slivkins, A.A. (2011). Contextual bandits with similarity information. JMLR: Workshop and Conference Proceedings 19 679–701
2011
Closest in time.