Fetching the paper…
Reading the bibliography…
We study a multi-armed bandit problem where the rewards exhibit regime switching.
Weighted sums of certain dependent random variables
K. Azuma · 1967
Earlier work this paper cites.
Arbitrary state markovian decision processes
S. M. Ross · 1968
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 2002
Earlier work this paper cites.
Discretized approximations for pomdp with average cost
H. Yu and D. P. Bertsekas · 2004
Earlier work this paper cites.
Lipschitz continuity of value functions in markovian decision processes
K. Hinderer · 2005
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
P. Auer and R. Ortner · 2007
Earlier work this paper cites.
Hidden Markov models in finance , volume 4
R. S. Mamon and R. J. Elliott · 2007
Earlier work this paper cites.
Adapting to a changing environment: the brownian restless bandits
A. Slivkins and E. Upfal · 2008
Earlier work this paper cites.
A structured multiarmed bandit problem and the greedy policy
A. J. Mersereau, P. Rusmevichientong, and J. N. Tsitsiklis · 2009
Earlier work this paper cites.
Trend following trading under a regime switching model
M. Dai, Q. Zhang, and Q. J. Zhu · 2010
Earlier work this paper cites.
Approximation algorithms for restless bandit problems
S. Guha, K. Munagala, and P. Shi · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
A. Garivier and E. Moulines · 2011
Earlier work this paper cites.
A method of moments for mixture models and hidden markov models
A. Anandkumar, D. Hsu, and S. M. Kakade · 2012
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Online regret bounds for undiscounted continuous reinforcement learning
R. Ortner and D. Ryabko · 2012
Cited alongside, same era.
Tensor decompositions for learning latent variable models
A. Anandkumar, R. Ge, D. Hsu, S. M. Kakade, and M. Telgarsky · 2014
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
O. Besbes, Y. Gur, and A. Zeevi · 2014
Cited alongside, same era.
Bounded regret for finite-armed structured bandits
T. Lattimore and R. Munos · 2014
Cited alongside, same era.
Latent bandits
O. A. Maillard and S. Mannor · 2014
Cited alongside, same era.
Learning to optimize under non-stationarity
W. C. Cheung, D. Simchi-Levi, and R. Zhu · 2018
Later among the works it cites.
Correlated multi-armed bandits with a latent random source
S. Gupta, G. Joshi, and O. Yağan · 2018
Later among the works it cites.
Deep variational reinforcement learning for pomdps
M. Igl, L. Zintgraf, T. A. Le, F. Wood, and S. Whiteson · 2018
Later among the works it cites.
Bandit algorithms
T. Lattimore and C. Szepesvári · 2018
Later among the works it cites.
J. Qian, R. Fruit, M. Pirotta, and A. Lazaric · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regret bounds for restless markov bandits
R. Ortner, D. Ryabko, P. Auer, and R. Munos · 2014
Cited alongside, same era.
Improved regret bounds for undiscounted continuous reinforcement learning
K. Lakshmanan, R. Ortner, and D. Ryabko · 2015
Cited alongside, same era.
Reinforcement learning of POMDPs using spectral methods
K. Azizzadenesheli, A. Lazaric, and A. Anandkumar · 2016
Cited alongside, same era.
Partially observed Markov decision processes
V. Krishnamurthy · 2016
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
S. Agrawal and R. Jia · 2017
Cited alongside, same era.
Consistent estimation of the filtering and marginal smoothing distributions in nonparametric hidden markov models
Y. De Castro, E. Gassiat, and S. Le Corff · 2017
Cited alongside, same era.
Adaptively tracking the best bandit arm with an unknown number of distribution changes
P. Auer, P. Gajane, and R. Ortner · 2019
Later among the works it cites.
Optimal exploration-exploitation in a multi-armed-bandit problem with non-stationary rewards
O. Besbes, Y. Gur, and A. Zeevi · 2019
Later among the works it cites.
A survey on practical applications of multi-armed and contextual bandits
D. Bouneffouf and I. Rish · 2019
Later among the works it cites.
The complexity of pomdps with long-run average objectives
K. Chatterjee, R. Saona, and B. Ziliotto · 2019
Later among the works it cites.
Multi-armed bandits for correlated markovian environments with smoothed reward feedback
T. Fiez, S. Sekar, and L. J. Ratliff · 2019
Later among the works it cites.
Consistent order estimation for nonparametric hidden markov models
L. Lehéricy · 2019
Later among the works it cites.
Introduction to multi-armed bandits
A. Slivkins · 2019
Later among the works it cites.
Learning and optimization with seasonal patterns
N. Chen, C. Wang, and L. Wang · 2020
Closest in time.
Approximate relative value learning for average-reward continuous state mdps
H. Sharma, M. Jafarnia-Jahromi, and R. Jain · 2020
Closest in time.