Fetching the paper…
Reading the bibliography…
In many sequential decision-making problems, the individuals are split into several batches and the decision-maker is only allowed to change her policy at the end of batches.
Deep neural linear bandits: Overcoming catastrophic forgetting through likelihood matching
Zahavy, T · 1901
Earlier work this paper cites.
A survey on practical applications of multi-armed and contextual bandits
Bouneffouf, D · 1904
Earlier work this paper cites.
Batched multi-armed bandits with optimal regret
Esfandiari, H · 1910
Earlier work this paper cites.
Neural contextual bandits with UCB-based exploration
Zhou, D · 1911
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P · 2002
Earlier work this paper cites.
Double explore-then-commit: Asymptotic optimality and beyond
Jin, T · 2002
Earlier work this paper cites.
Sequential batch learning in finite-action linear contextual bandits
Han, Y · 2004
Earlier work this paper cites.
The epoch-greedy algorithm for contextual multi-armed bandits
Langford, J · 2007
Earlier work this paper cites.
Linear bandits with limited adaptivity and learning distributional optimal design
Ruan, Y · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V · 2008
Earlier work this paper cites.
Crowdsourcing user studies with mechanical turk
Kittur, A · 2008
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, N · 2009
Cited alongside, same era.
Ucb revisited: Improved regret bounds for the stochastic multi-armed bandit problem
Auer, P · 2010
Cited alongside, same era.
Parametric bandits: The generalized linear case
Filippi, S · 2010
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Li, L · 2010
Cited alongside, same era.
Zhang, W · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Finite-time analysis of kernelised contextual bandits
Valko, M · 2013
Later among the works it cites.
Bandits with switching costs: T 2/3 regret
Dekel, O · 2014
Later among the works it cites.
Batched bandit problems
Perchet, V · 2016
Later among the works it cites.
Provably optimal algorithms for generalized linear contextual bandits
Li, L · 2017
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An empirical evaluation of thompson sampling
Chapelle, O · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Chu, W · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S · 2012
Cited alongside, same era.
Neural contextual bandits with deep representation and shallow exploration
Xu, P · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S · 2013
Cited alongside, same era.
Online learning with switching costs and other adaptive adversaries
Cesa-Bianchi, N · 2013
Cited alongside, same era.
Later among the works it cites.
Riquelme, C · 2018
Later among the works it cites.
Batched multi-armed bandits problem
Gao, Z · 2019
Later among the works it cites.
Batch policy learning under constraints
Le, H · 2019
Later among the works it cites.
Phase transitions and cyclic phenomena in bandits with switching constraints
Simchi-Levi, D · 2019
Later among the works it cites.
Introduction to multi-armed bandits
Slivkins, A · 2019
Later among the works it cites.
Bandit Algorithms
Lattimore, T · 2020
Later among the works it cites.