Fetching the paper…
Reading the bibliography…
We consider reinforcement learning in changing Markov Decision Processes where both the state-transition probabilities and the reward functions may vary over time.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Optimal adaptive policies for markov decision processes
Apostolos N. Burnetas and Michael N. Katehakis · 1997
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Experts in a Markov decision process
Eyal Even-dar, Sham M Kakade, and Yishay Mansour · 2005
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Cited alongside, same era.
The robustness-performance tradeoff in Markov decision processes
Huan Xu and Shie Mannor · 2006
Cited alongside, same era.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L. Bartlett and Ambuj Tewari · 2009
Cited alongside, same era.
Online learning in Markov decision processes with arbitrarily changing rewards and transitions
Jia Yuan Yu and Shie Mannor · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Arbitrarily modulated Markov decision processes
Jia Yuan Yu and Shie Mannor
Cited in the paper.
On upper-confidence bound policies for switching bandit problems
Aurélien Garivier and Eric Moulines · 2011
Later among the works it cites.
Online learning in Markov decision processes with adversarially chosen transition probability distributions
Yasin Abbasi, Peter L Bartlett, Varun Kanade, Yevgeny Seldin, and Csaba Szepesvari · 2013
Later among the works it cites.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Later among the works it cites.
Online learning in Markov decision processes with changing cost sequences
T Dick, András György, and Csaba Szepesvári · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…