2019

Neural Policy Gradient Methods: Global Optimality and Rates of Convergence

Wang, Lingxiao, Cai, Qi, Yang, Zhuoran et al.

Understand

Policy gradient methods with actor-critic schemes demonstrate tremendous empirical successes, especially when the actors and critics are parameterized by neural networks.

  • However, it remains less clear whether such "neural" policy gradient methods converge to globally optimal policies and whether they even converge at all.
  • We answer both the questions affirmatively in the overparameterized regime.
  • In detail, we prove that neural natural policy gradient converges to a globally optimal policy at a sublinear rate.

Reading the bibliography…