Fetching the paper…
Reading the bibliography…
As an important type of reinforcement learning algorithms, actor-critic (AC) and natural actor-critic (NAC) algorithms are often executed in two ways for finding optimal policies.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D. (2019) · 1906
Earlier work this paper cites.
Global convergence of policy gradient methods to (almost) locally optimal policies
Zhang, K., Koppel, A., Zhu, H., and Başar, T. (2019) · 1906
Earlier work this paper cites.
Optimality and approximation with policy gradient methods in Markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2019) · 1908
Earlier work this paper cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Shani, L., Efroni, Y., and Mannor, S. (2019) · 1909
Earlier work this paper cites.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L., Cai, Q., Yang, Z., and Wang, Z. (2019) · 1909
Earlier work this paper cites.
Kumar, H., Koppel, A., and Ribeiro, A. (2019) · 1910
Earlier work this paper cites.
A tale of two-timescale reinforcement learning with the tightest finite-time bound
Dalal, G., Szorenyi, B., and Thoppe, G. (2019) · 1911
Earlier work this paper cites.
Non-asymptotic analysis of biased stochastic approximation scheme
Karimi, B., Miasojedow, B., Moulines, E., and Wai, H.-T. (2019) · 1974
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Stochastic approximation with two time scales
Borkar, V. S. (1997) · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
Actor-critic–type learning algorithms for Markov decision processes
Konda, V. R. and Borkar, V. S. (1999) · 1999
Earlier work this paper cites.
Reinforcement learning in continuous time and space
Doya, K. (2000) · 2000
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L. (2001) · 2001
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2002) · 2002
Cited alongside, same era.
Finite time analysis of linear two-timescale stochastic approximation with Markovian noise
Kaledin, M., Moulines, E., Naumov, A., Tadic, V., and Wai, H.-T. (2020) · 2002
Cited alongside, same era.
Actor-critic algorithms
Konda, V. (2002) · 2002
Cited alongside, same era.
Non-asymptotic convergence of Adam-type reinforcement learning algorithms under markovian sampling
Xiong, H., Xu, T., Liang, Y., and Zhang, W. (2020) · 2002
Cited alongside, same era.
Almost sure convergence of two time-scale stochastic approximation algorithms
Tadic, V. B. (2004) · 2004
Cited alongside, same era.
Improving sample complexity bounds for actor-critic algorithms
Adaptive batch size for safe policy gradients
Papini, M., Pirotta, M., and Restelli, M. (2017) · 2017
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J., Russo, D., and Singal, R. (2018) · 2018
Later among the works it cites.
Concentration bounds for two time scale stochastic approximation
Borkar, V. S. and Pattathil, S. (2018) · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Fazel, M., Ge, R., Kakade, S. M., and Mesbahi, M. (2018) · 2018
Later among the works it cites.
Derivative-free methods for policy optimization: guarantees for linear quadratic systems
Malik, D., Pananjady, A., Bhatia, K., Khamaru, K., Bartlett, P. L., and Wainwright, M. J. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xu, T., Wang, Z., and Liang, Y. (2020b) · 2004
Cited alongside, same era.
Convergence rate and averaging of nonlinear two-time-scale stochastic approximation algorithms
Mokkadem, A. and Pelletier, M. (2006) · 2006
Cited alongside, same era.
Incremental natural actor-critic algorithms
Bhatnagar, S., Ghavamzadeh, M., Lee, M., and Sutton, R. S. (2008) · 2008
Cited alongside, same era.
Natural actor-critic
Peters, J. and Schaal, S. (2008) · 2008
Cited alongside, same era.
Natural actor–critic algorithms
Bhatnagar, S., Sutton, R. S., Ghavamzadeh, M., and Lee, M. (2009) · 2009
Cited alongside, same era.
Stochastic approximation: a dynamical systems viewpoint
Borkar, V. S. (2009) · 2009
Cited alongside, same era.
An actor–critic algorithm with function approximation for discounted cost constrained Markov decision processes
Bhatnagar, S. (2010) · 2010
Cited alongside, same era.
Stochastic variance-reduced policy gradient
Papini, M., Binaghi, D., Canonaco, G., Pirotta, M., and Restelli, M. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Tu, S. and Recht, B. (2018) · 2018
Later among the works it cites.
Finite-time performance bounds and adaptive learning rate selection for two time-scale reinforcement learning
Gupta, H., Srikant, R., and Ying, L. (2019) · 2019
Later among the works it cites.
Neural proximal/trust region policy optimization attains globally optimal policy
Liu, B., Cai, Q., Yang, Z., and Wang, Z. (2019) · 2019
Later among the works it cites.
On the finite-time convergence of actor-critic algorithm
Qiu, S., Yang, Z., Ye, J., and Wang, Z. (2019) · 2019
Later among the works it cites.
Hessian aided policy gradient
Shen, Z., Ribeiro, A., Hassani, H., Qian, H., and Mi, C. (2019) · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation and TD learning
Srikant, R. and Ying, L. (2019) · 2019
Later among the works it cites.
Provably global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost
Yang, Z., Chen, Y., Hong, M., and Wang, Z. (2019) · 2019
Later among the works it cites.
Finite-sample analysis for SARSA with linear function approximation
Zou, S., Xu, T., and Liang, Y. (2019) · 2019
Later among the works it cites.