Fetching the paper…
Reading the bibliography…
We consider the problem of exploration-exploitation in communicating Markov Decision Processes.
Exploration-exploitation tradeoff using variance estimates in multi-armed bandits
Audibert, J.-Y., Munos, R., and Szepesvári, C. (2009) · 1902
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, D. A. (1975) · 1975
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Influence and variance of a markov chain: Application to adaptive discretization in optimal control
Munos, R. and Moore, A. (1999) · 1999
Earlier work this paper cites.
Tuning bandit algorithms in stochastic environments
Audibert, J.-Y., Munos, R., and Szepesvári, C. (2007) · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Cited alongside, same era.
Pac bounds for discounted mdps
Lattimore, T. and Hutter, M. (2012) · 2012
Cited alongside, same era.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J. (2013) · 2013
Cited alongside, same era.
Near-optimal pac bounds for discounted mdps
Lattimore, T. and Hutter, M. (2014) · 2014
Cited alongside, same era.
How hard is my mdp?” the distribution-norm to the rescue”
Maillard, O.-A., Mann, T. A., and Mannor, S. (2014) · 2014
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Later among the works it cites.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Fruit, R., Pirotta, M., Lazaric, A., and Ortner, R. (2018) · 2018
Later among the works it cites.
Bandit algorithms
Lattimore, T. and Szepesvári, C. (2018) · 2018
Later among the works it cites.
Variance-aware regret bounds for undiscounted reinforcement learning in mdps
Talebi, M. S. and Maillard, O. (2018) · 2018
Later among the works it cites.
Exploration-exploitation dilemma in Reinforcement Learning under various form of prior knowledge
Fruit, R. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…