Fetching the paper…
Reading the bibliography…
We present simple and efficient algorithms for the batched stochastic multi-armed bandit and batched stochastic linear bandit problems.
Introduction to multi-armed bandits
Aleksandrs Slivkins · 1935
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Batch mode active learning and its application to medical image classification
Steven CH Hoi, Rong Jin, Jianke Zhu, and Michael R Lyu · 2006
Earlier work this paper cites.
A learning approach for interactive marketing to a customer segment
Dimitris Bertsimas and Adam J Mersereau · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
Crowdsourcing user studies with mechanical turk
Aniket Kittur, Ed H Chi, and Bongwon Suh · 2008
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
Jean-Yves Audibert and Sébastien Bubeck · 2009
Earlier work this paper cites.
UCB revisited: Improved regret bounds for the stochastic multi-armed bandit problem
Peter Auer and Ronald Ortner · 2010
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
Aurélien Garivier and Eric Moulines · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Towards minimax policies for online linear optimization with bandit feedback
Sébastien Bubeck, Nicoló Cesa-Bianchi, and Sham M. Kakade · 2012
Cited alongside, same era.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Cited alongside, same era.
Online learning with switching costs and other adaptive adversaries
Nicolò Cesa-Bianchi, Ofer Dekel, and Ohad Shamir · 2013
Cited alongside, same era.
Near-optimal batch mode active learning and adaptive submodular optimization
Yuxin Chen and Andreas Krause · 2013
Cited alongside, same era.
Learning with limited rounds of adaptivity: Coin tossing, multi-armed bandits, and ranking from pairwise comparisons
Arpit Agarwal, Shivani Agarwal, Sepehr Assadi, and Sanjeev Khanna · 2017
Later among the works it cites.
The adaptive complexity of maximizing a submodular function
Eric Balkanski and Yaron Singer · 2018
Later among the works it cites.
High-dimensional probability: An introduction, with applications in data science , volume 47 of Cambridge Series in Statistical and Probabilistic Mathematics
Roman Vershynin · 2018
Later among the works it cites.
Stochastic submodular cover with limited adaptivity
Arpit Agarwal, Sepehr Assadi, and Sanjeev Khanna · 2019
Closest in time.
Delay and cooperation in nonstochastic bandits
Nicolò Cesa-Bianchi, Claudio Gentile, and Yishay Mansour · 2019
Closest in time.
Unconstrained submodular maximization with constant adaptive complexity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Parallel Gaussian process optimization with upper confidence bound and pure exploration
Emile Contal, David Buffoni, Alexandre Robicquet, and Nicolas Vayatis · 2013
Cited alongside, same era.
Parallelizing exploration-exploitation tradeoffs in Gaussian process bandit optimization
Thomas Desautels, Andreas Krause, and Joel W Burdick · 2014
Cited alongside, same era.
Batched bandit problems
Vianney Perchet, Philippe Rigollet, Sylvain Chassang, and Erik Snowberg · 2015
Cited alongside, same era.
Scalable Bayesian optimization using deep neural networks
Jasper Snoek, Oren Rippel, Kevin Swersky, Ryan Kiros, Nadathur Satish, Narayanan Sundaram, Mostofa Patwary, Mr Prabhat, and Ryan Adams · 2015
Cited alongside, same era.
Top arm identification in multi-armed bandits with batch arm pulls
Kwang-Sung Jun, Kevin Jamieson, Robert Nowak, and Xiaojin Zhu · 2016
Cited alongside, same era.
Batched Gaussian process bandit optimization via determinantal point processes
Tarun Kathuria, Amit Deshpande, and Pushmeet Kohli · 2016
Cited alongside, same era.
Lin Chen, Moran Feldman, and Amin Karbasi · 2019
Closest in time.
Adaptivity in adaptive submodularity
Hossein Esfandiari, Amin Karbasi, and Vahab Mirrokni · 2019
Closest in time.
Submodular maximization with optimal approximation, adaptivity and query complexity
Matthew Fahrbach, Vahab Mirrokni, and Morteza Zadimoghaddam · 2019
Closest in time.
Batched multi-armed bandits problem
Zijun Gao, Yanjun Han, Zhimei Ren, and Zhengqing Zhou · 2019
Closest in time.
Batch policy learning under constraints
Hoang M Le, Cameron Voloshin, and Yisong Yue · 2019
Closest in time.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48 of Cambridge Series in Statistical and Probabilistic Mathematics
Martin J. Wainwright · 2019
Closest in time.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Closest in time.