Fetching the paper…
Reading the bibliography…
Stochastic Lipschitz bandit algorithms balance exploration and exploitation, and have been used for a variety of important task domains.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson. 1933 · 1933
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
John C Gittins. 1979 · 1979
Earlier work this paper cites.
Classification and regression trees
Leo Breiman, Jerome Friedman, Charles J Stone, and Richard A Olshen. 1984 · 1984
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins. 1985 · 1985
Earlier work this paper cites.
Incremental induction of decision trees
Paul E Utgoff. 1989 · 1989
Earlier work this paper cites.
Introduction to reinforcement learning . Vol. 135
Richard S Sutton and Andrew G Barto. 1998 · 1998
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer. 2002 · 2002
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer. 2002 · 2002
Earlier work this paper cites.
Improved rates for the stochastic continuum-armed bandit problem. In International Conference on Computational Learning Theory . Springer
Peter Auer, Ronald Ortner, and Csaba Szepesvári. 2007 · 2007
Earlier work this paper cites.
Stochastic Linear Optimization under Bandit Feedback.. In Annual Conference on Learning Theory . 355–366
Varsha Dani, Thomas P Hayes, and Sham M Kakade. 2008 · 2008
Earlier work this paper cites.
Multi-armed bandits in metric spaces. In Proceedings of the fortieth annual ACM symposium on Theory of computing . ACM, 681–690
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal. 2008 · 2008
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th international conference on World Wide Web . ACM, 661–670
Lihong Li, Wei Chu, John Langford, and Robert E Schapire. 2010 · 2010
Cited alongside, same era.
Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design. In Proceedings of the 27th International Conference on Machine Learning . Omnipress, Haifa, Israel, 1015–1022
Niranjan Srinivas, Andreas Krause, Sham Kakade, and Matthias Seeger. 2010 · 2010
Cited alongside, same era.
Improved Algorithms for Linear Stochastic Bandits
Yasin Abbasi-yadkori, Dávid Pál, and Csaba Szepesvári. 2011 · 2011
Cited alongside, same era.
X-armed bandits
Sébastien Bubeck, Rémi Munos, Gilles Stoltz, and Csaba Szepesvári. 2011 · 2011
Exponential Regret Bounds for Gaussian Process Bandits with Deterministic Observations. In Proceedings of International Conference on Machine Learning
Nando de Freitas, Alex Smola, and Masrour Zoghi. 2012 · 2012
Later among the works it cites.
Thompson Sampling for Contextual Bandits with Linear Payoffs (Proceedings of Machine Learning Research, Vol. 28) . PMLR, Atlanta, Georgia, USA, 127–135
Shipra Agrawal and Navin Goyal. 2013 · 2013
Later among the works it cites.
Gaussian process optimization with mutual information. In Proceedings of International Conference on Machine Learning . 253–261
Emile Contal, Vianney Perchet, and Nicolas Vayatis. 2014 · 2014
Later among the works it cites.
Lipschitz bandits: Regret lower bound and optimal algorithms. In Annual Conference on Learning Theory . 975–999
Stefan Magureanu, Richard Combes, and Alexandre Proutiere. 2014 · 2014
Later among the works it cites.
Bayesopt: A Bayesian optimization library for nonlinear optimization, experimental design and bandits
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scikit-learn: Machine Learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Cited alongside, same era.
A variant of Azuma’s inequality for martingales with sub-Gaussian tails
Ohad Shamir. 2011 · 2011
Cited alongside, same era.
Analysis of Thompson Sampling for the Multi-armed Bandit Problem (Proceedings of Machine Learning Research, Vol. 23) . JMLR Workshop and Conference Proceedings, Edinburgh, Scotland, 39.1–39.26
Shipra Agrawal and Navin Goyal. 2012 · 2012
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolo Cesa-Bianchi. 2012 · 2012
Cited alongside, same era.
Ruben Martinez-Cantin. 2014 · 2014
Later among the works it cites.
Contextual bandits with similarity information
Aleksandrs Slivkins. 2014 · 2014
Later among the works it cites.
Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Later among the works it cites.
Hyperband: Bandit-based configuration evaluation for hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar. 2016 · 2016
Later among the works it cites.
Learning to control in metric space with optimal regret. In 57th Annual Allerton Conference on Communication, Control, and Computing (Allerton) . IEEE, 726–733
Chengzhuo Ni, Lin F Yang, and Mengdi Wang. 2019 · 2019
Closest in time.