2019

From Importance Sampling to Doubly Robust Policy Gradient

Huang, Jiawei, Jiang, Nan

Understand

We show that on-policy policy gradient (PG) and its variance reduction variants can be derived by taking finite difference of function evaluations supplied by estimators from the importance sampling (IS) family for off-policy evaluation (OPE).

  • Starting from the doubly robust (DR) estimator (Jiang & Li, 2016), we provide a simple derivation of a very general and flexible form of PG, which subsumes the state-of-the-art variance reduction technique (Cheng et al., 2019) as its special case and immediately hints at further variance reduction opportunities overlooked by existing literature.
  • We analyze the variance of the new DR-PG estimator, compare it to existing methods as well as the Cramer-Rao lower bound of policy gradient, and empirically show its effectiveness.

Built on

  • Simple statistical gradient-following algorithms for connectionist reinforcement learning

    Williams, R. J · 1992

    Earlier work this paper cites.

  • Eligibility traces for off-policy policy evaluation

    Precup, D · 2000

    Earlier work this paper cites.

  • Policy gradient methods for reinforcement learning with function approximation

    Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000

    Earlier work this paper cites.

  • Infinite-horizon policy-gradient estimation

    Baxter, J. and Bartlett, P. L · 2001

    Earlier work this paper cites.

  • Variance reduction techniques for gradient estimates in reinforcement learning

    Greensmith, E., Bartlett, P. L., and Baxter, J · 2004

    Earlier work this paper cites.

  • A Theory of Cramer-Rao Bounds for Constrained Parametric Models

    Moore, T. J · 2010

    Earlier work this paper cites.

Similar

  • On a connection between importance sampling and the likelihood ratio policy gradient

    Tang, J. and Abbeel, P · 2010

    Cited alongside, same era.

  • The dependence of effective planning horizon on model accuracy

    Jiang, N., Kulesza, A., Singh, S., and Lewis, R · 2015

    Cited alongside, same era.

  • Openai gym

    Original

    Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016

    Cited alongside, same era.

  • Doubly robust off-policy value evaluation for reinforcement learning

    Jiang, N. and Li, L · 2016

    Cited alongside, same era.

  • Data-efficient off-policy policy evaluation for reinforcement learning

    Thomas, P. and Brunskill, E · 2016

    Cited alongside, same era.

  • Q-prop: Sample-efficient policy gradient with an off-policy critic

    Gu, S., Lillicrap, T. P., Ghahramani, Z., Turner, R. E., and Levine, S · 2017

    Cited alongside, same era.

Then

  • Backpropagation through the void: Optimizing control variates for black-box gradient estimation

    Grathwohl, W., Choi, D., Wu, Y., Roeder, G., and Duvenaud, D · 2018

    Later among the works it cites.

  • DART: dynamic animation and robotics toolkit

    Lee, J., Grey, M. X., Ha, S., Kunz, T., Jain, S., Ye, Y., Srinivasa, S. S., Stilman, M., and Liu, C. K · 2018

    Later among the works it cites.

  • Action-dependent control variates for policy optimization via stein identity

    Liu, H., Feng, Y., Mao, Y., Zhou, D., Peng, J., and Liu, Q · 2018

    Later among the works it cites.

  • The mirage of action-dependent baselines in reinforcement learning

    Tucker, G., Bhupatiraju, S., Gu, S., Turner, R., Ghahramani, Z., and Levine, S · 2018

    Later among the works it cites.

  • Variance reduction for policy gradient with action-dependent factorized baselines

    Wu, C., Rajeswaran, A., Duan, Y., Kumar, V., Bayen, A. M., Kakade, S., Mordatch, I., and Abbeel, P · 2018

    Later among the works it cites.

  • Trajectory-wise control variates for variance reduction in policy gradient methods

    Cheng, C.-A., Yan, X., and Boots., B · 2019

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…