Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel algorithm named STOchastic Recursive Momentum for Policy Gradient (STORM-PG), which operates a SARAH-type stochastic recursive variance-reduced policy gradient in an exponential moving average fashion.
An improved convergence analysis of stochastic variance-reduced policy gradient
Xu, P., Gao, F., and Gu, Q · 1905
Earlier work this paper cites.
Sample efficient policy gradient methods with recursive variance reduction
Xu, P., Gao, F., and Gu, Q · 1909
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
On-line Q-learning using connectionist systems , volume 37
Rummery, G. A. and Niranjan, M · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Tesauro, G · 1995
Earlier work this paper cites.
Reinforcement learning for continuous stochastic control problems
Munos, R. and Bourgine, P · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S., Barto, A. G., Bach, F., et al · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L · 2001
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Cited alongside, same era.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Later among the works it cites.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, L. M., Liu, J., Scheinberg, K., and Takáč, M · 2017
Later among the works it cites.
Stochastic variance reduction for policy gradient estimation
Xu, T., Liu, Q., and Peng, J · 2017
Later among the works it cites.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Fang, C., Li, C. J., Lin, Z., and Zhang, T · 2018
Later among the works it cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S., Shammah, S., and Shashua, A · 2016
Cited alongside, same era.
Stochastic variance-reduced policy gradient
Papini, M., Binaghi, D., Canonaco, G., Pirotta, M., and Restelli, M · 2018
Later among the works it cites.
Lower bounds for non-convex stochastic optimization
Arjevani, Y., Carmon, Y., Duchi, J. C., Foster, D. J., Srebro, N., and Woodworth, B · 2019
Later among the works it cites.
Momentum-based variance reduction in non-convex sgd
Cutkosky, A. and Orabona, F · 2019
Later among the works it cites.