Fetching the paper…
Reading the bibliography…
We study the multi-armed bandit problem with subgaussian rewards.
Batched multi-armed bandits with optimal regret
Esfandiari, H · 1910
Earlier work this paper cites.
Étude critique de la notion de collectif
Ville, J · 1939
Earlier work this paper cites.
Parallelism in comparison problems
Valiant, L. G · 1975
Earlier work this paper cites.
Parallel sorting
Bollobás, B · 1983
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L · 1985
Earlier work this paper cites.
Deterministic selection in o (loglog n) parallel time
Ajtai, M · 1986
Earlier work this paper cites.
Sorting, approximate sorting, and searching in rounds
Alon, N · 1988
Earlier work this paper cites.
Computing with noisy information
Feige, U · 1994
Earlier work this paper cites.
Sequential choice from several populations
Katehakis, M. N · 1995
Earlier work this paper cites.
Sequential batch learning in finite-action linear contextual bandits
Han, Y · 2004
Earlier work this paper cites.
A learning approach for interactive marketing to a customer segment
Bertsimas, D · 2007
Earlier work this paper cites.
Linear bandits with limited adaptivity and learning distributional optimal design
Ruan, Y · 2007
Cited alongside, same era.
Minimax policies for adversarial and stochastic bandits
Audibert, J.-Y · 2009
Cited alongside, same era.
Economic analysis of simulation selection problems
Chick, S. E · 2009
Cited alongside, same era.
The kl-ucb algorithm for bounded stochastic bandits and beyond
Garivier, A · 2011
Cited alongside, same era.
Online learning with switching costs and other adaptive adversaries
Cesa-Bianchi, N · 2013
Cited alongside, same era.
Thompson sampling for 1-dimensional exponential family bandits
Korda, N · 2013
Cited alongside, same era.
Learning with limited rounds of adaptivity: Coin tossing, multi-armed bandits, and ranking from pairwise comparisons
Agarwal, A · 2017
Later among the works it cites.
Near-optimal regret bounds for thompson sampling
Agrawal, S · 2017
Later among the works it cites.
A minimax and asymptotically optimal algorithm for stochastic bandits
Ménard, P · 2017
Later among the works it cites.
Minimax bounds on stochastic batched convex optimization
Duchi, J · 2018
Later among the works it cites.
On bayesian index policies for sequential resource allocation
Kaufmann, E · 2018
Later among the works it cites.
Refining the confidence level for optimistic bandit strategies
Lattimore, T · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Parallel algorithms for select and partition with noisy comparisons
Braverman, M · 2016
Cited alongside, same era.
Anytime optimal algorithms in stochastic multi-armed bandits
Degenne, R · 2016
Cited alongside, same era.
Optimal best arm identification with fixed confidence
Garivier, A · 2016
Cited alongside, same era.
On explore-then-commit strategies
Garivier, A · 2016
Cited alongside, same era.
Batched bandit problems
Perchet, V · 2016
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Auer, P
Cited in the paper.
Probability: theory and examples
Durrett, R · 2019
Later among the works it cites.
Batched multi-armed bandits problem
Gao, Z · 2019
Later among the works it cites.
Efficient pure exploration in adaptive round model
Jin, T · 2019
Later among the works it cites.
Collaborative learning with limited interaction: Tight bounds for distributed exploration in multi-armed bandits
Tao, C · 2019
Later among the works it cites.
Bandit algorithms
Lattimore, T · 2020
Closest in time.