Fetching the paper…
Reading the bibliography…
Markov Decision Processes are classically solved using Value Iteration and Policy Iteration algorithms.
The theory of dynamic programming
R. Bellman · 1954
Earlier work this paper cites.
On the convergence of policy iteration in stationary dynamic programming
M. L. Puterman and S. L. Brumelle · 1979
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
S. Nemirovsky, Arkadiĭ and D. B. Yudin · 1983
Earlier work this paper cites.
Real applications of markov decision processes
D. J. White · 1985
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
M. L. Puterman · 1995
Earlier work this paper cites.
Elements of information theory
T. M. Cover · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al · 1999
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Natural actor-critic
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Natural actor–critic algorithms
S. Bhatnagar, R. S. Sutton, M. Ghavamzadeh, and M. Lee · 2009
Earlier work this paper cites.
A generalized natural actor-critic algorithm
T. Morimura, E. Uchibe, J. Yoshimoto, and K. Doya · 2009
Earlier work this paper cites.
Dynamic policy programming
M. G. Azar, V. Gómez, and H. J. Kappen · 2012
Cited alongside, same era.
Projected natural actor-critic
P. S. Thomas, W. C. Dabney, S. Giguere, and S. Mahadevan · 2013
Cited alongside, same era.
The information geometry of mirror descent
G. Raskutti and S. Mukherjee · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Improved and generalized upper bounds on the complexity of policy iteration
B. Scherrer · 2016
Cited alongside, same era.
First-order methods in optimization
A. Beck · 2017
Cited alongside, same era.
A note on the linear convergence of policy gradient methods
J. Bhandari and D. Russo · 2020
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
S. Cen, C. Cheng, Y. Chen, Y. Wei, and Y. Chi · 2020
Later among the works it cites.
Statistics, Computation, and Adaptation in High Dimensions
A. P. Martin · 2020
Later among the works it cites.
On the global convergence rates of softmax policy gradient methods
J. Mei, C. Xiao, C. Szepesvari, and D. Schuurmans · 2020
Later among the works it cites.
Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs
L. Shani, Y. Efroni, and S. Mannor · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2019
Cited alongside, same era.
A theory of regularized markov decision processes
M. Geist, B. Scherrer, and O. Pietquin · 2019
Cited alongside, same era.
On the finite-time convergence of actor-critic algorithm
S. Qiu, Z. Yang, J. Ye, and Z. Wang · 2019
Cited alongside, same era.
Adaptive sampling for best policy identification in markov decision processes
A. Al Marjani and A. Proutiere · 2020
Cited alongside, same era.
A finite time analysis of two time-scale actor critic methods
Y. Wu, W. Zhang, P. Xu, and Q. Gu · 2020
Later among the works it cites.
Improving sample complexity bounds for actor-critic algorithms
T. Xu, Z. Wang, and Y. Liang · 2020
Later among the works it cites.
Variational policy gradient method for reinforcement learning with general utilities
J. Zhang, A. Koppel, A. S. Bedi, C. Szepesvari, and M. Wang · 2020
Later among the works it cites.
Finite-sample analysis of off-policy natural actor-critic algorithm
S. Khodadadian, Z. Chen, and S. T. Maguluri · 2021
Closest in time.
Finite Sample Analysis of Two-Time-Scale Natural Actor-Critic Algorithm
S. Khodadadian, T. T. Doan, S. T. Maguluri, and J. Romberg · 2021
Closest in time.