Fetching the paper…
Reading the bibliography…
We consider the problem of learning the optimal policy for infinite-horizon Markov decision processes (MDPs).
Problem Complexity and Method Efficiency in Optimization
Nemirovski, A. and Yudin, D. (1983) · 1983
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
Convex Optimization
Boyd, S. and Vandenberghe, L. (2004) · 2004
Earlier work this paper cites.
Dynamic Programming and Optimal Control
Bertsekas, D. (2005) · 2005
Earlier work this paper cites.
Primal-dual subgradient methods for convex problems
Nesterov, Y. (2009) · 2005
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L. (2008) · 2005
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Bartlett, P. L. and Tewari, A. (2009) · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A. (2009) · 2009
Earlier work this paper cites.
Nemirovski’s inequalities revisited
Dümbgen, L., Geer, S., Veraar, M., and Wellner, J. (2010) · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Cited alongside, same era.
Algorithms for Reinforcement Learning
Szepesvari, C. (2010) · 2010
Cited alongside, same era.
On the Sample Complexity of Reinforcement Learning with a Generative Model
Azar, M. G., Munos, R., and Kappen, B. (2012) · 2012
Cited alongside, same era.
Validation analysis of mirror descent stochastic approximation method
Lan, G., Nemirovski, A., and Shapiro, A. (2012) · 2012
Cited alongside, same era.
Concentration Inequalities: A Nonasymptotic Theory of Independence
Boucheron, S., Lugosi, G., and Massart, P. (2013) · 2013
Cited alongside, same era.
Primal-dual subgradient method for huge-scale linear conic problems
Nesterov, Y. and Shpirko, S. (2014) · 2014
Cited alongside, same era.
Mirror Descent and Convex Optimization Problems with Non-smooth Inequality Constraints
Bayandina, A., Dvurechensky, P., Gasnikov, A., Stonyakin, F., and Titov, A. (2018) · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Adaptive mirror descent algorithms for convex and strongly convex optimization problems with functional constraints
Stonyakin, F., Alkousa, M., Stepanov, A., and Tytov, A. (2019) · 2019
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
Agarwal, A., Kakade, S., and Yang, L. F. (2020) · 2020
Later among the works it cites.
Efficiently solving MDPs with stochastic mirror descent
Jin, Y. and Sidford, A. (2020) · 2020
Later among the works it cites.
Algorithms for stochastic optimization with expectation constraints
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On lower complexity bounds for large-scale smooth convex optimization
Guzmán, C. and Nemirovski, A. (2015) · 2015
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, S. and Jia, R. (2017) · 2017
Cited alongside, same era.
Primal-dual π \pi learning: Sample complexity and sublinear run time for ergodic markov decision problems
Wang, M. (2017) · 2017
Cited alongside, same era.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
Sidford, A., Wang, M., Wu, X., Yang, L. F., and Ye, Y. (2018a)
Cited in the paper.
Variance reduced value iteration and faster algorithms for solving markov decision processes
Sidford, A., Wang, M., Wu, X., and Ye, Y. (2018b)
Cited in the paper.
Lan, G. and Zhou, Z. (2020) · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Li, G., Wei, Y., Chi, Y., Gu, Y., and Chen, Y. (2020) · 2020
Later among the works it cites.
Towards tight bounds on the sample complexity of average-reward mdps
Jin, Y. and Sidford, A. (2021) · 2021
Closest in time.