Fetching the paper…
Reading the bibliography…
Despite their success, policy gradient methods suffer from high variance of the gradient estimate, which can result in unsatisfactory sample complexity.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L · 2001
Earlier work this paper cites.
Analysis and improvement of policy gradient estimation
Zhao, T., Hachiya, H., Niu, G., and Sugiyama, M · 2011
Earlier work this paper cites.
Stochastic gradient descent tricks
Bottou, L · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Learning complex neural network policies with trajectory optimization
Levine, S. and Koltun, V · 2014
Earlier work this paper cites.
Adding gradient noise improves learning for very deep networks
Neelakantan, A., Vilnis, L., Le, Q. V., Sutskever, I., Kaiser, L., Kurach, K., and Martens, J · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Allen-Zhu, Z · 2017
Earlier work this paper cites.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, L. M., Liu, J., Scheinberg, K., and Takáč, M · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Stochastic variance-reduced policy gradient
Papini, M., Binaghi, D., Canonaco, G., Pirotta, M., and Restelli, M · 2018
Cited alongside, same era.
Variance reduced value iteration and faster algorithms for solving markov decision processes
Sidford, A., Wang, M., Wu, X., and Ye, Y · 2018
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P., Ying, C., and Le, Q. V · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
An improved convergence analysis of stochastic variance-reduced policy gradient
Xu, P., Gao, F., and Gu, Q · 2019
Later among the works it cites.
Towards understanding the importance of noise in training neural networks
Zhou, M., Liu, T., Li, Y., Lin, D., Zhou, E., and Zhao, T · 2019
Later among the works it cites.
On the promise of the stochastic generalized Gauss-Newton method for training DNNs
Gargiani, M., Zanelli, A., Diehl, M., and Hutter, F · 2020
Later among the works it cites.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2020
Later among the works it cites.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
On the theory of policy gradient methods: optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Gaurav, M · 2019
Cited alongside, same era.
Understanding the impact of entropy on policy optimization
Ahmed, Z., Le Roux, N., Norouzi, M., and Schuurmans, D · 2019
Cited alongside, same era.
Momentum-based variance reduction in non-convex SGD
Cutkosky, A. and Orabona, F · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
Kovalev, D., Horváth, S., and Richtárik, P · 2020
Later among the works it cites.
Stochastic recursive momentum for policy gradient methods
Yuan, H., Lian, X., Liu, J., and Zhou, Y · 2020
Later among the works it cites.
On the linear convergence of policy gradient methods for finite MDPs
Bhandari, J. and Russo, D · 2021
Later among the works it cites.
Page: A simple and optimal probabilistic gradient estimator for nonconvex optimization
Li, Z., Bao, H., Zhang, X., and Richtarik · 2021
Later among the works it cites.
Sample efficient policy gradient methods with recursive variance reduction
Xu, P., Gao, F., and Gu, Q · 2021
Later among the works it cites.
On the convergence and sample efficiency of variance-reduced policy gradient method
Zhang, J., Ni, C., Yu, Z., Szepesvari, C., and Wang, M · 2021
Later among the works it cites.