Fetching the paper…
Reading the bibliography…
Upper Confidence Bound (UCB) algorithms are a widely-used class of sequential algorithms for the $K$-armed bandit problem.
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvári, Exploration-exploitation tradeoff using variance estimates in multi-armed bandits , Theoret. Comput. Sci. 410
1902
Earlier work this paper cites.
Victor H. de la Peña, Michael J. Klass, and Tze Leung Lai, Self-normalized processes: exponential inequalities, moment bounds and iterated logarithm laws , Ann. Probab. 32
1933
Earlier work this paper cites.
William R Thompson, On the likelihood that one unknown probability exceeds another in view of the evidence of two samples , Biometrika 25
1933
Earlier work this paper cites.
Herbert Robbins, Some aspects of the sequential design of experiments , Bull. Amer. Math. Soc. 58
1952
Earlier work this paper cites.
John S. White, The limiting distribution of the serial correlation coefficient in the explosive case , Ann. Math. Statist. 29
1958
Earlier work this paper cites.
by same author, The limiting distribution of the serial correlation coefficient in the explosive case. II , Ann. Math. Statist. 30
1959
Earlier work this paper cites.
Aryeh Dvoretzky, Asymptotic normality for sums of dependent random variables , Proceedings of the Sixth Berkeley Symposium on Mathematical Statistics and Probability (Univ. California, Berkeley, Calif., 1970/1971), Vol. II: Probability theory, Univ. California Press, Berkeley, CA, 1972, pp. 513–535
1972
Earlier work this paper cites.
David A. Dickey and Wayne A. Fuller, Distribution of the estimators for autoregressive time series with a unit root , J. Amer. Statist. Assoc. 74
1979
Earlier work this paper cites.
P. Hall and C. C. Heyde, Martingale limit theory and its application , Probability and Mathematical Statistics, Academic Press, Inc., New York-London, 1980
1980
Earlier work this paper cites.
Tze Leung Lai and Ching Zong Wei, Least squares estimates in stochastic regression models with applications to identification and control of dynamic systems , Ann. Statist. 10
1982
Earlier work this paper cites.
Tze Leung Lai and Herbert Robbins, Asymptotically efficient adaptive allocation rules , Adv. in Appl. Math. 6
1985
Earlier work this paper cites.
David Siegmund, Sequential analysis , Springer Series in Statistics, Springer-Verlag, New York, 1985, Tests and confidence intervals
1985
Earlier work this paper cites.
by same author, Boundary crossing probabilities and statistical applications , Ann. Statist. 14
1986
Earlier work this paper cites.
Tze Leung Lai, Adaptive treatment allocation and the multi-armed bandit problem , Ann. Statist. 15
1987
Earlier work this paper cites.
K. Joos, Nonuniform convergence rates in the central limit theorem for martingales , Studia Sci. Math. Hungar. 28
1993
Earlier work this paper cites.
by same author, Gambling in a rigged casino: the adversarial multi-armed bandit problem , 36th Annual Symposium on Foundations of Computer Science (Milwaukee, WI, 1995), IEEE Comput. Soc. Press, Los Alamitos, CA, 1995, pp. 322–331
1995
Earlier work this paper cites.
Rajeev Agrawal, Sample mean based index policies with O ( log n ) O(\log n) regret for the multi-armed bandit problem , Adv. in Appl. Probab. 27
1995
Earlier work this paper cites.
Michael N. Katehakis and Herbert Robbins, Sequential choice from several populations , Proc. Nat. Acad. Sci. U.S.A. 92
1995
Earlier work this paper cites.
Rajendra Bhatia, Matrix analysis , Graduate Texts in Mathematics, vol. 169, Springer-Verlag, New York, 1997
1997
Earlier work this paper cites.
Víctor H. de la Peña and Evarist Giné, Decoupling , Probability and its Applications (New York), Springer-Verlag, New York, 1999, From dependence to independence, Randomly stopped processes. U U -statistics and processes. Martingales and beyond
1999
Earlier work this paper cites.
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer, Finite-time analysis of the multiarmed bandit problem , Machine Learning 47
2002
Earlier work this paper cites.
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire, The nonstochastic multiarmed bandit problem , SIAM J. Comput. 32
2002
Cited alongside, same era.
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvári, Tuning bandit algorithms in stochastic environments , International Conference on Algorithmic Learning Theory, Springer, 2007, pp. 150–165
2007
Cited alongside, same era.
Pierre-Arnaud Coquelin and Rémi Munos, Bandit algorithms for tree search , Proceedings of the Twenty-Third Conference on Uncertainty in Artificial Intelligence (Arlington, Virginia, USA), UAI’07, AUAI Press, 2007, p. 67–74
2007
Cited alongside, same era.
Jean-Yves Audibert and Sébastien Bubeck, Minimax policies for adversarial and stochastic bandits , COLT, 2009, pp. 217–226
2009
Cited alongside, same era.
Aleksandrs Slivkins, Introduction to multi-armed bandits , Foundations and Trends® in Machine Learning 12
2019
Later among the works it cites.
2019
Later among the works it cites.
Tor Lattimore and Csaba Szepesvári, Bandit algorithms , Cambridge University Press, 2020
2020
Later among the works it cites.
Kelly Zhang, Lucas Janson, and Susan Murphy, Inference for batched bandits , Advances in Neural Information Processing Systems 33
2020
Later among the works it cites.
Aurélien Bibaut, Maria Dimakopoulou, Nathan Kallus, Antoine Chambaz, and Mark van Der Laan, Post-contextual-bandit inference , Advances in Neural Information Processing Systems 34
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2009
Cited alongside, same era.
Junya Honda and Akimichi Takemura, An asymptotically optimal bandit algorithm for bounded support models. , COLT, Citeseer, 2010, pp. 67–79
2010
Cited alongside, same era.
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári, Improved algorithms for linear stochastic bandits , Advances in Neural Information Processing Systems 24
2011
Cited alongside, same era.
Sébastien Bubeck and Nicolo Cesa-Bianchi, Regret analysis of stochastic and nonstochastic multi-armed bandit problems , Foundations and Trends® in Machine Learning 5
2012
Cited alongside, same era.
Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier, On bayesian upper confidence bounds for bandit problems , Artificial Intelligence and Statistics, PMLR, 2012, pp. 592–600
2012
Cited alongside, same era.
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart, Concentration inequalities: A nonasymptotic theory of independence , Oxford University Press, Oxford, 2013
2013
Cited alongside, same era.
Olivier Cappé, Aurélien Garivier, Odalric-Ambrym Maillard, Rémi Munos, and Gilles Stoltz, Kullback-Leibler upper confidence bounds for optimal sequential allocation , Ann. Statist. 41
2013
Cited alongside, same era.
Jean-Christophe Mourrat, On the rate of convergence in the martingale central limit theorem , Bernoulli 19
2013
Cited alongside, same era.
Later among the works it cites.
Vitor Hadad, David A. Hirshberg, Ruohan Zhan, Stefan Wager, and Susan Athey, Confidence intervals for policy evaluation in adaptive experiments , Proc. Natl. Acad. Sci. USA 118
2021
Later among the works it cites.
2021
Later among the works it cites.
Anand Kalvit and Assaf Zeevi, A closer look at the worst-case behavior of multi-armed bandit algorithms , Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Ruohan Zhan, Vitor Hadad, David A Hirshberg, and Susan Athey, Off-policy evaluation via adaptive weighting with data from contextual bandits , Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, 2021, pp. 2125–2135
2021
Later among the works it cites.
Aurélien Garivier, Hédi Hadiji, Pierre Ménard, and Gilles Stoltz, KL-UCB-switch: optimal regret bounds for stochastic bandits from both a distribution-dependent and a distribution-free viewpoints , J. Mach. Learn. Res. 23
2022
Later among the works it cites.
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J. Tibshirani, Surprises in high-dimensional ridgeless least squares interpolation , Ann. Statist. 50
2022
Later among the works it cites.
2023
Later among the works it cites.
Yash Deshpande, Adel Javanmard, and Mohammad Mehrabi, Online debiasing for adaptively collected high-dimensional data with applications to time series analysis , J. Amer. Statist. Assoc. 118
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Licong Lin, Mufang Ying, Suvrojit Ghosh, Koulik Khamaru, and Cun-Hui Zhang, Statistical limits of adaptive linear models: low-dimensional estimation and inference , Advances in Neural Information Processing Systems 36
2024
Closest in time.
2024
Closest in time.
Mufang Ying, Koulik Khamaru, and Cun-Hui Zhang, Adaptive linear estimating equations , Advances in Neural Information Processing Systems 36
2024
Closest in time.