Fetching the paper…
Reading the bibliography…
We consider the restless Markov bandit problem, in which the state of each arm evolves according to a Markov process independently of the learner's actions.
Bandit processes and dynamic allocation indices
Gittins, J.C.: · 1979
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T.L., Robbins, H.: · 1985
Earlier work this paper cites.
Asymptotically efficient allocation rules for the multiarmed bandit problem with multiple plays, part II: Markovian rewards
Anantharam, V., Varaiya, P., Walrand, J.: · 1987
Earlier work this paper cites.
Restless bandits: Activity allocation in a changing world
Whittle, P.: · 1988
Earlier work this paper cites.
Threshold limits for cover times
Aldous, D.: · 1991
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., Schapire, R.E.: · 2002
Cited alongside, same era.
Finite-time analysis of the multi-armed bandit problem
Auer, P., Cesa-Bianchi, N., Fischer, P.: · 2002
Cited alongside, same era.
Markov chains and mixing times
Levin, D.A., Peres, Y., Wilmer, E.L.: · 2006
Cited alongside, same era.
Pseudometrics for state aggregation in average reward Markov decision processes
Ortner, R.: · 2007
Cited alongside, same era.
A survey on spectrum management in cognitive radio networks
Akyildiz, I.F., Lee, W.Y.L.W.Y., Vuran, M.C., Mohanty, S.: · 2008
Cited alongside, same era.
Reversible Markov Chains and Random Walks on Graphs
Aldous, D.J., Fill, J.:
Cited in the paper.
Minimax policies for adversarial and stochastic bandits
Audibert, J.Y., Bubeck, S.: · 2009
Later among the works it cites.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Bartlett, P.L., Tewari, A.: · 2009
Later among the works it cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., Auer, P.: · 2010
Later among the works it cites.
Adaptive learning of uncontrolled restless bandits with logarithmic regret
Tekin, C., Liu, M.: · 2011
Later among the works it cites.
Optimally sensing a single channel without prior information: The tiling algorithm and regret bounds
Filippi, S., Cappé and, O., Garivier, A.: · 2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…