Understand
We show that on-policy policy gradient (PG) and its variance reduction variants can be derived by taking finite difference of function evaluations supplied by estimators from the importance sampling (IS) family for off-policy evaluation (OPE).
- Starting from the doubly robust (DR) estimator (Jiang & Li, 2016), we provide a simple derivation of a very general and flexible form of PG, which subsumes the state-of-the-art variance reduction technique (Cheng et al., 2019) as its special case and immediately hints at further variance reduction opportunities overlooked by existing literature.
- We analyze the variance of the new DR-PG estimator, compare it to existing methods as well as the Cramer-Rao lower bound of policy gradient, and empirically show its effectiveness.
Built on
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L · 2001
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, E., Bartlett, P. L., and Baxter, J · 2004
Earlier work this paper cites.
A Theory of Cramer-Rao Bounds for Constrained Parametric Models
Moore, T. J · 2010
Earlier work this paper cites.
Similar
On a connection between importance sampling and the likelihood ratio policy gradient
Tang, J. and Abbeel, P · 2010
Cited alongside, same era.
The dependence of effective planning horizon on model accuracy
Jiang, N., Kulesza, A., Singh, S., and Lewis, R · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Gu, S., Lillicrap, T. P., Ghahramani, Z., Turner, R. E., and Levine, S · 2017
Cited alongside, same era.
Then
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Grathwohl, W., Choi, D., Wu, Y., Roeder, G., and Duvenaud, D · 2018
Later among the works it cites.
DART: dynamic animation and robotics toolkit
Lee, J., Grey, M. X., Ha, S., Kunz, T., Jain, S., Ye, Y., Srinivasa, S. S., Stilman, M., and Liu, C. K · 2018
Later among the works it cites.
Action-dependent control variates for policy optimization via stein identity
Liu, H., Feng, Y., Mao, Y., Zhou, D., Peng, J., and Liu, Q · 2018
Later among the works it cites.
The mirage of action-dependent baselines in reinforcement learning
Tucker, G., Bhupatiraju, S., Gu, S., Turner, R., Ghahramani, Z., and Levine, S · 2018
Later among the works it cites.
Variance reduction for policy gradient with action-dependent factorized baselines
Wu, C., Rajeswaran, A., Duan, Y., Kumar, V., Bayen, A. M., Kakade, S., Mordatch, I., and Abbeel, P · 2018
Later among the works it cites.
Trajectory-wise control variates for variance reduction in policy gradient methods
Cheng, C.-A., Yan, X., and Boots., B · 2019
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…