Fetching the paper…
Reading the bibliography…
We prove new upper and lower bounds for sample complexity of finding an $\epsilon$-optimal policy of an infinite-horizon average-reward Markov decision process (MDP) given access to a generative model.
Average reward reinforcement learning: Foundations, algorithms, and empirical results
S. Mahadevan · 1996
Earlier work this paper cites.
B. Yu · 1997
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade et al · 2003
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
P. Ortner and R. Auer · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
P. L. Bartlett and A. Tewari · 2012
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
M. G. Azar, R. Munos, and H. J. Kappen · 2013
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 2014
Cited alongside, same era.
Faster algorithms for computing the stationary distribution, simulating random walks, and more
M. B. Cohen, J. Kelner, J. Peebles, R. Peng, A. Sidford, and A. Vladu · 2016
Cited alongside, same era.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
R. Fruit, M. Pirotta, A. Lazaric, and R. Ortner · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
How does an approximate model help in reinforcement learning?
Variance-reduced q-learning is minimax optimal
M. J. Wainwright · 2019
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
A. Agarwal, S. Kakade, and L. F. Yang · 2020
Later among the works it cites.
Efficiently solving mdps with stochastic mirror descent
Y. Jin and A. Sidford · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
G. Li, Y. Wei, Y. Chi, Y. Gu, and Y. Chen · 2020
Later among the works it cites.
Regret bounds for reinforcement learning via markov chain concentration
R. Ortner · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Feng, W. Yin, and L. F. Yang · 2019
Cited alongside, same era.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
A. Sidford, M. Wang, X. Wu, L. Yang, and Y. Ye
Cited in the paper.
Variance reduced value iteration and faster algorithms for solving markov decision processes
A. Sidford, M. Wang, X. Wu, and Y. Ye
Cited in the paper.
M. Wang
Cited in the paper.
M. Wang
Cited in the paper.
D. Tiapkin, F. Stonyakin, and A. Gasnikov · 2021
Closest in time.