Fetching the paper…
Reading the bibliography…
Actor-critic (AC) methods have exhibited great empirical success compared with other reinforcement learning algorithms, where the actor uses the policy gradient to improve the learning policy and the critic uses temporal difference learning to estimate the policy gradient.
Global convergence of policy gradient methods to (almost) locally optimal policies
Kaiqing Zhang, Alec Koppel, Hao Zhu, and Tamer Başar · 1906
Earlier work this paper cites.
Provably convergent two-timescale off-policy actor-critic with function approximation
Shangtong Zhang, Bo Liu, Hengshuai Yao, and Shimon Whiteson · 1911
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
A. G. Barto, R. S. Sutton, and C. W. Anderson · 1983
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Stochastic approximation with two time scales
Vivek S Borkar · 1997
Earlier work this paper cites.
The actor-critic algorithm as multi-time-scale stochastic approximation
Vivek S Borkar and Vijaymohan R Konda · 1997
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Asymptotic properties of two time-scale stochastic approximation algorithms with constant step sizes
Vladislav B Tadic and Sean P Meyn · 2003
Earlier work this paper cites.
Convergence rate of linear two-time-scale stochastic approximation
Vijay R Konda, John N Tsitsiklis, et al · 2004
Earlier work this paper cites.
Convergence and divergence in standard and averaging reinforcement learning
Marco A Wiering · 2004
Earlier work this paper cites.
Sensitivity and convergence of uniformly ergodic markov chains
A Yu Mitrophanov · 2005
Earlier work this paper cites.
Non-asymptotic convergence analysis of two time-scale (natural) actor-critic algorithms
Tengyu Xu, Zhe Wang, and Yingbin Liang · 2005
Cited alongside, same era.
Natural actor–critic algorithms
Shalabh Bhatnagar, Richard S Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2009
Cited alongside, same era.
A convergent online single time scale actor critic algorithm
Dotan Di Castro and Ron Meir · 2010
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2016
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2019
Later among the works it cites.
Performance of q-learning with linear function approximation: Stability and finite-time analysis
Zaiwei Chen, Sheng Zhang, Thinh T Doan, Siva Theja Maguluri, and John-Paul Clarke · 2019
Later among the works it cites.
Finite-time performance bounds and adaptive learning rate selection for two time-scale reinforcement learning
Harsh Gupta, R Srikant, and Lei Ying · 2019
Later among the works it cites.
Characterizing the exact behaviors of temporal difference learning algorithms using markov jump linear system theory
Bin Hu and Usman Syed · 2019
Later among the works it cites.
Harshat Kumar, Alec Koppel, and Alejandro Ribeiro · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2016
Cited alongside, same era.
Gal Dalal, Balazs Szorenyi, Gugan Thoppe, and Shie Mannor · 2017
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Lectures on convex optimization , volume 137
Yurii Nesterov · 2018
Cited alongside, same era.
Stochastic variance-reduced policy gradient
Matteo Papini, Damiano Binaghi, Giuseppe Canonaco, Matteo Pirotta, and Marcello Restelli · 2018
Cited alongside, same era.
Later among the works it cites.
On the finite-time convergence of actor-critic algorithm
Shuang Qiu, Zhuoran Yang, Jieping Ye, and Zhaoran Wang · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation andtd learning
R Srikant and Lei Ying · 2019
Later among the works it cites.
A finite-time analysis of q-learning with neural network function approximation
Pan Xu and Quanquan Gu · 2019
Later among the works it cites.
On the global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost
Zhuoran Yang, Yongxin Chen, Mingyi Hong, and Zhaoran Wang · 2019
Later among the works it cites.
Finite-sample analysis for sarsa with linear function approximation
Shaofeng Zou, Tengyu Xu, and Yingbin Liang · 2019
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2020
Closest in time.
Finite time analysis of linear two-timescale stochastic approximation with markovian noise
Maxim Kaledin, Eric Moulines, Alexey Naumov, Vladislav Tadic, and Hoi-To Wai · 2020
Closest in time.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2020
Closest in time.