Fetching the paper…
Reading the bibliography…
We study a multi-armed bandit problem in a dynamic environment where arm rewards evolve in a correlated fashion according to a Markov chain.
Introduction to Mathematical Probability
J. Uspensky · 1937
Earlier work this paper cites.
Dynamic programming, princeton university pres
R. Bellman · 1957
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
W. Hoeffding · 1963
Earlier work this paper cites.
Weighted sums of certain dependent random variables
K. Azuma · 1967
Earlier work this paper cites.
Choices, values, and frames
D. Kahneman and A. Tversky · 1984
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1985
Earlier work this paper cites.
Asymptotically efficient allocation rules for the multiarmed bandit problem with multiple plays–Part II: Markovian rewards
V. Anantharam, P. Varaiya, and J. Walrand · 1987
Earlier work this paper cites.
Eigenvalue bounds on convergence to stationarity for nonreversible Markov chains, with an application to the exclusion process
J. A. Fill · 1991
Earlier work this paper cites.
Contract theory
P. Bolton, M. Dewatripont, et al · 2005
Earlier work this paper cites.
Real Analysis
G. Folland · 2007
Earlier work this paper cites.
Trueskill™: a Bayesian skill rating system
R. Herbrich, T. Minka, and T. Graepel · 2007
Earlier work this paper cites.
Mortal multi-armed bandits
D. Chakrabarti, R. Kumar, F. Radlinski, and E. Upfal · 2008
Earlier work this paper cites.
An analysis of reinforcement learning with function approximation
F. S. Melo, S. P. Meyn, and M. I. Ribeiro · 2008
Cited alongside, same era.
Subdominant eigenvalues for stochastic matrices with given column sums
S. Kirkland · 2009
Cited alongside, same era.
The theory of incentives: the principal-agent model
J.-J. Laffont and D. Martimort · 2009
Cited alongside, same era.
Web-scale Bayesian click-through rate prediction for sponsored search advertising in Microsoft’s Bing search engine
T. Graepel, J. Q. Candela, T. Borchert, and R. Herbrich · 2010
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Regret bounds for reinforcement learning with policy advice
M. G. Azar, A. Lazaric, and E. Brunskill · 2013
Later among the works it cites.
A tutorial on linear function approximators for dynamic programming and reinforcement learning
A. Geramifard, T. J. Walsh, S. Tellex, G. Chowdhary, N. Roy, and J. P. How · 2013
Later among the works it cites.
Regret bounds for restless markov bandits
R. Ortner, D. Ryabko, P. Auer, and R. Munos · 2014
Later among the works it cites.
Efficient crowdsourcing of unknown experts using bounded multi-armed bandits
L. Tran-Thanh, S. Stein, A. Rogers, and N. R. Jennings · 2014
Later among the works it cites.
Adaptive contract design for crowdsourcing markets: Bandit algorithms for repeated principal-agent problems
C.-J. Ho, A. Slivkins, and J. W. Vaughan · 2016
Later among the works it cites.
Collaborative filtering bandits
S. Li, A. Karatzoglou, and C. Gentile · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Online algorithms for the multi-armed bandit problem with Markovian rewards
C. Tekin and M. Liu · 2010
Cited alongside, same era.
On the combinatorial multi-armed bandit problem with Markovian rewards
Y. Gai, B. Krishnamachari, and M. Liu · 2011
Cited alongside, same era.
On upper-confidence bound policies for switching bandit problems
A. Garivier and E. Moulines · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Online bandit learning against an adaptive adversary: from regret to policy regret
O. Dekel, A. Tewari, and R. Arora · 2012
Cited alongside, same era.
Online learning of rested and restless bandits
C. Tekin and M. Liu · 2012
Cited alongside, same era.
Later among the works it cites.
Discrepancy-based algorithms for non-stationary rested bandits
C. Cortes, G. DeSalvo, V. Kuznetsov, M. Mohri, and S. Yang · 2017
Later among the works it cites.
Rotting bandits
N. Levine, K. Crammer, and S. Mannor · 2017
Later among the works it cites.
An online learning approach to improving the quality of crowd-sourcing
Y. Liu and M. Liu · 2017
Later among the works it cites.
A multi-armed bandit approach for online expert selection in markov decision processes
E. Mazumdar, R. Dong, V. R. Royo, C. Tomlin, and S. S. Sastry · 2017
Later among the works it cites.
Recharging bandits
N. Immorlica and R. D. Kleinberg · 2018
Closest in time.
Bandit learning with positive externalities
V. Shah, J. Blanchet, and R. Johari · 2018
Closest in time.