Fetching the paper…
Reading the bibliography…
Despite the empirical success of the actor-critic algorithm, its theoretical understanding lags behind.
On the global convergence of imitation learning: A case for linear quadratic regulator
Cai, Q · 1901
Earlier work this paper cites.
The bias and moment matrix of the general k-class estimators of the parameters in simultaneous equations
Nagar, A. L · 1959
Earlier work this paper cites.
Gradient methods for minimizing functionals
Polyak, B. T · 1963
Earlier work this paper cites.
Optimal control theory: an introduction
Kirk, D. E · 1970
Earlier work this paper cites.
The moments of products of quadratic forms in normal variables
Magnus, J. R · 1978
Earlier work this paper cites.
Estimation of the mean of a multivariate normal distribution
Stein, C. M · 1981
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
New branch-and-bound rules for linear bilevel programming
Hansen, P · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Reinforcement learning applied to linear quadratic regulation
Bradtke, S. J · 1993
Earlier work this paper cites.
Adaptive linear quadratic control using policy iteration
Bradtke, S. J · 1994
Earlier work this paper cites.
Mathematical programs with equilibrium constraints
Luo, Z.-Q · 1996
Earlier work this paper cites.
Zhou, K · 1996
Earlier work this paper cites.
Stochastic approximation with two time scales
Borkar, V. S · 1997
Earlier work this paper cites.
The actor-critic algorithm as multi-time-scale stochastic approximation
Borkar, V. S · 1997
Earlier work this paper cites.
The policy iteration algorithm for average reward markov decision processes with general state space
Meyn, S. P · 1997
Earlier work this paper cites.
Primal-dual interior-point methods for semidefinite programming: convergence rates, stability and numerical results
Alizadeh, F · 1998
Earlier work this paper cites.
Dynamic noncooperative game theory
Basar, T · 1999
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J · 2001
Cited alongside, same era.
Foundations of bilevel programming
Dempe, S · 2002
Cited alongside, same era.
Stochastic approximation and recursive algorithms and applications
Kushner, H · 2003
Cited alongside, same era.
Optimal control: linear quadratic methods
Anderson, B. D · 2007
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A · 2009
Cited alongside, same era.
A review of stochastic algorithms with continuous value function approximation and some new approximate policy iteration algorithms for multidimensional continuous applications
Powell, W. B · 2011
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Dean, S · 2017
Later among the works it cites.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Islam, R · 2017
Later among the works it cites.
Certifiable distributional robustness with principled adversarial training
Sinha, A · 2017
Later among the works it cites.
Least-squares temporal difference learning for the linear quadratic regulator
Tu, S · 2017
Later among the works it cites.
Finite sample analysis of the gtd policy evaluation algorithms in markov setting
Wang, Y · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamic programming and optimal control, Vol. II, 4th Edition
Bertsekas, D. P · 2012
Cited alongside, same era.
A survey of actor-critic reinforcement learning: Standard and natural policy gradients
Grondman, I · 2012
Cited alongside, same era.
Practical bilevel optimization: algorithms and applications
Bard, J. F · 2013
Cited alongside, same era.
Matrix analysis
Horn, R. A · 2013
Cited alongside, same era.
Hanson-wright inequality and sub-gaussian concentration
Rudelson, M · 2013
Cited alongside, same era.
Policy evaluation with temporal differences: A survey and comparison
Dann, C · 2014
Cited alongside, same era.
Chen, X · 2018
Later among the works it cites.
Du, S. S · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Fazel, M · 2018
Later among the works it cites.
Gradient descent learns linear dynamical systems
Hardt, M · 2018
Later among the works it cites.
Lin, Q · 2018
Later among the works it cites.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
Malik, D · 2018
Later among the works it cites.
Non-convex min-max optimization: Provable algorithms and applications in machine learning
Rafique, H · 2018
Later among the works it cites.
A tour of reinforcement learning: The view from continuous control
Recht, B · 2018
Later among the works it cites.
Solving non-convex non-concave min-max games under polyak-lojasiewicz condition
Sanjabi, M · 2018
Later among the works it cites.
Learning without mixing: Towards a sharp analysis of linear system identification
Simchowitz, M · 2018
Later among the works it cites.
A review on bilevel optimization: from classical to evolutionary approaches and applications
Sinha, A · 2018
Later among the works it cites.
Tu, S · 2018
Later among the works it cites.
Convergent reinforcement learning with function approximation: A bilevel optimization perspective
Yang, Z · 2018
Later among the works it cites.
Understand the dynamics of GANs via primal-dual optimization. https://openreview.net/forum?id=rylIy3R9K7
Lu, S · 2019
Closest in time.