Fetching the paper…
Reading the bibliography…
Consider the problem of sampling sequentially from a finite number of $N \geq 2$ populations, specified by random variables $X^i_k$, $ i = 1,\ldots , N,$ and $k = 1, 2, \ldots$; where $X^i_k$ denotes the outcome from population $i$ the $k^{th}$ time it is sampled.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Bandit processes and dynamic allocation indices (with discussion)
John C. Gittins · 1979
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
On the Gittins index for multiarmed bandits
Richard R Weber · 1992
Earlier work this paper cites.
On sequencing two types of tasks on a single processor under incomplete information
Apostolos N Burnetas and Michael N Katehakis · 1993
Earlier work this paper cites.
Sequential choice from several populations
Michael N Katehakis and Herbert Robbins · 1995
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Asymptotic Bayes analysis for the finite-horizon one-armed-bandit problem
Apostolos N Burnetas and Michael N Katehakis · 2003
Earlier work this paper cites.
Cooperative Control: Models, Applications, and Algorithms
Sergiy Butenko, Panos M Pardalos, and Robert Murphey · 2003
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible mdps
Ambuj Tewari and Peter L Bartlett · 2008
Cited alongside, same era.
Exploration–exploitation tradeoff using variance estimates in multi-armed bandits
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvári · 2009
Cited alongside, same era.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L Bartlett and Ambuj Tewari · 2009
Cited alongside, same era.
Multi-armed bandit based policies for cognitive radio’s decision making issues
Wassim Jouini, Damien Ernst, Christophe Moy, and Jacques Palicot · 2009
Cited alongside, same era.
Ucb revisited: Improved regret bounds for the stochastic multi-armed bandit problem
Peter Auer and Ronald Ortner · 2010
Cited alongside, same era.
Optimism in reinforcement learning based on kullbackleibler divergence
Inducing partially observable Markov decision processes
Michael L Littman · 2012
Later among the works it cites.
Approximately optimal adaptive learning in opportunistic spectrum access
Cem Tekin and Mingyan Liu · 2012
Later among the works it cites.
Kullback–leibler upper confidence bounds for optimal sequential allocation
Olivier Cappé, Aurélien Garivier, Odalric-Ambrym Maillard, Rémi Munos, and Gilles Stoltz · 2013
Later among the works it cites.
Asymptotically optimal Bayesian sequential change detection and identification rules
Savas Dayanik, Warren B Powell, and Kazutoshi Yamazaki · 2013
Later among the works it cites.
The multi-armed bandit, with constraints
Eric V Denardo, Eugene A Feinberg, and Uriel G Rothblum · 2013
Later among the works it cites.
Optimality of Thompson sampling for Gaussian bandits depends on priors
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sarah Filippi, Olivier Cappé, and Aurélien Garivier · 2010
Cited alongside, same era.
An asymptotically optimal bandit algorithm for bounded support models
Junya Honda and Akimichi Takemura · 2010
Cited alongside, same era.
Multi-armed Bandit Allocation Indices
John C. Gittins, Kevin Glazebrook, and Richard R. Weber · 2011
Cited alongside, same era.
An asymptotically optimal policy for finite support models in the multiarmed bandit problem
Junya Honda and Akimichi Takemura · 2011
Cited alongside, same era.
The best of both worlds: Stochastic and adversarial bandits
Sébastien Bubeck and Aleksandrs Slivkins · 2012
Cited alongside, same era.
On large deviations properties of sequential allocation problems
Apostolos N Burnetas and Michael N Katehakis
Cited in the paper.
Optimal adaptive policies for sequential allocation problems
Apostolos N Burnetas and Michael N Katehakis
Cited in the paper.
Junya Honda and Akimichi Takemura · 2013
Later among the works it cites.
Convergence of value iterations for total-cost mdps and pomdps with general state and action sets
Eugene A Feinberg, Pavlo O Kasyanov, and Michael Z Zgurovsky · 2014
Later among the works it cites.
On minimax optimal offline policy evaluation
Lihong Li, Remi Munos, and Csaba Szepesvari · 2014
Later among the works it cites.
Near-optimal reinforcement learning in factored mdps
Ian Osband and Benjamin Van Roy · 2014
Later among the works it cites.
Analyse de stratégies Bayésiennes et fréquentistes pour l’allocation séquentielle de ressources
Emilie Kaufmann · 2015
Closest in time.