Fetching the paper…
Reading the bibliography…
Natural policy gradient (NPG) methods are among the most widely used policy optimization algorithms in contemporary reinforcement learning.
Non-asymptotic analysis of biased stochastic approximation scheme
Karimi, B., Miasojedow, B., Moulines, É., and Wai, H.-T. (2019) · 1902
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D. (2019) · 1906
Earlier work this paper cites.
Global convergence of policy gradient methods to (almost) locally optimal policies
Zhang, K., Koppel, A., Zhu, H., and Başar, T. (2019b) · 1906
Earlier work this paper cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized MDPs
Shani, L., Efroni, Y., and Mannor, S. (2019) · 1909
Earlier work this paper cites.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L., Cai, Q., Yang, Z., and Wang, Z. (2019) · 1909
Earlier work this paper cites.
Zhang, K., Hu, B., and Basar, T. (2019a) · 1910
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z. (2019) · 1912
Earlier work this paper cites.
Mohammadi, H., Zare, A., Soltanolkotabi, M., and Jovanović, M. R. (2019) · 1912
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
On the theory of dynamic programming
Bellman, R. (1952) · 1952
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovsky, A. S. and Yudin, D. B. (1983) · 1983
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J. (1991) · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
Elements of information theory
Cover, T. M. (1999) · 1999
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Convergence guarantees of policy optimization methods for Markovian jump linear systems
Jansch-Porto, J. P., Hu, B., and Dullerud, G. (2020) · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Cited alongside, same era.
A natural policy gradient
Kakade, S. M. (2002) · 2002
Cited alongside, same era.
Leverage the average: an analysis of regularization in RL
Vieillard, N., Kozuno, T., Scherrer, B., Pietquin, O., Munos, R., and Geist, M. (2020) · 2003
Cited alongside, same era.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Li, G., Wei, Y., Chi, Y., Gu, Y., and Chen, Y. (2020a) · 2005
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvari, C., and Schuurmans, D. (2020) · 2005
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Later among the works it cites.
Dynamic programming and optimal control (4th edition)
Bertsekas, D. P. (2017) · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. (2017) · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2017) · 2017
Later among the works it cites.
A unified view of entropy-regularized Markov decision processes
Neu, G., Jonsson, A., and Gómez, V. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Non-asymptotic convergence analysis of two time-scale (natural) actor-critic algorithms
Xu, T., Wang, Z., and Liang, Y. (2020) · 2005
Cited alongside, same era.
Sample complexity of asynchronous Q-learning: Sharper analysis and variance reduction
Li, G., Wei, Y., Chi, Y., Gu, Y., and Chen, Y. (2020b) · 2006
Cited alongside, same era.
A note on the linear convergence of policy gradient methods
Bhandari, J. and Russo, D. (2020) · 2007
Cited alongside, same era.
Natural actor-critic
Peters, J. and Schaal, S. (2008) · 2008
Cited alongside, same era.
Natural actor-critic algorithms
Bhatnagar, S., Sutton, R. S., Ghavamzadeh, M., and Lee, M. (2009) · 2009
Cited alongside, same era.
Online Markov decision processes
Even-Dar, E., Kakade, S. M., and Mansour, Y. (2009) · 2009
Cited alongside, same era.
Primal-dual subgradient methods for convex problems
Nesterov, Y. (2009) · 2009
Cited alongside, same era.
SBEED: Convergent reinforcement learning with nonlinear function approximation
Dai, B., Shaw, A., Li, L., Xiao, L., He, N., Liu, Z., Chen, J., and Song, L. (2018) · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Fazel, M., Ge, R., Kakade, S., and Mesbahi, M. (2018) · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
Reinforcement learning: Theory and algorithms
Agarwal, A., Jiang, N., and Kakade, S. M. (2019) · 2019
Later among the works it cites.
Understanding the impact of entropy on policy optimization
Ahmed, Z., Le Roux, N., Norouzi, M., and Schuurmans, D. (2019) · 2019
Later among the works it cites.
A theory of regularized Markov decision processes
Geist, M., Scherrer, B., and Pietquin, O. (2019) · 2019
Later among the works it cites.
Planning in entropy-regularized markov decision processes and games
Grill, J.-B., Darwiche Domingues, O., Menard, P., Munos, R., and Valko, M. (2019) · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S., Singh, K., and Van Soest, A. (2019) · 2019
Later among the works it cites.
Neural trust region/proximal policy optimization attains globally optimal policy
Liu, B., Cai, Q., Yang, Z., and Wang, Z. (2019) · 2019
Later among the works it cites.
The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint
Tu, S. and Recht, B. (2019) · 2019
Later among the works it cites.
Maximum entropy Monte-Carlo planning
Xiao, C., Huang, R., Mei, J., Schuurmans, D., and Müller, M. (2019) · 2019
Later among the works it cites.
A finite time analysis of two time-scale actor critic methods
Wu, Y., Zhang, W., Xu, P., and Gu, Q. (2020) · 2020
Closest in time.