Fetching the paper…
Reading the bibliography…
We revisit the stochastic variance-reduced policy gradient (SVRPG) method proposed by Papini et al.
Stochastic nested variance reduced gradient descent for nonconvex optimization
Zhou, D., Xu, P., and Gu, Q · 1905
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
On measures of entropy and information
Rényi, A. et al · 1961
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L · 2001
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Weaver, L. and Tao, N · 2001
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2002
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, E., Bartlett, P. L., and Baxter, J · 2004
Earlier work this paper cites.
Monte Carlo strategies in scientific computing
Liu, J. S · 2008
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Peters, J. and Schaal, S · 2008
Earlier work this paper cites.
Learning bounds for importance weighting
Cortes, C., Mansour, Y., and Mohri, M · 2010
Earlier work this paper cites.
Parameter-exploring policy gradients
Sehnke, F., Osendorfer, C., Rückstieß, T., Graves, A., Peters, J., and Schmidhuber, J · 2010
Earlier work this paper cites.
Analysis and improvement of policy gradient estimation
Zhao, T., Hachiya, H., Niu, G., and Sugiyama, M · 2011
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Cited alongside, same era.
Adaptive step-size for policy gradient methods
Pirotta, M., Restelli, M., and Bascetta, L · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Cited alongside, same era.
A proximal stochastic gradient method with progressive variance reduction
Xiao, L. and Zhang, T · 2014
Cited alongside, same era.
Stopwasting my gradients: Practical svrg
Harikandeh, R., Ahmed, M. O., Virani, A., Schmidt, M., Konečnỳ, J., and Sallinen, S · 2015
Cited alongside, same era.
Learning contact-rich manipulation skills with guided policy search
Levine, S., Wagener, N., and Abbeel, P · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T. P., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Later among the works it cites.
Stochastic variance reduction methods for policy evaluation
Du, S. S., Chen, J., Li, L., Xiao, L., and Zhou, D · 2017
Later among the works it cites.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., and Levine, S · 2017
Later among the works it cites.
Non-convex finite-sum optimization via scsg methods
Lei, L., Ju, C., Chen, J., and Jordan, M. I · 2017
Later among the works it cites.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, L. M., Liu, J., Scheinberg, K., and Takáč, M · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P · 2015
Cited alongside, same era.
Variance reduction for faster non-convex optimization
Allen-Zhu, Z. and Hazan, E · 2016
Cited alongside, same era.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Stochastic variance reduction for nonconvex optimization
Reddi, S. J., Hefny, A., Sra, S., Poczos, B., and Smola, A · 2016
Cited alongside, same era.
Adaptive batch size for safe policy gradients
Papini, M., Pirotta, M., and Restelli, M · 2017
Later among the works it cites.
Stochastic variance reduction for policy gradient estimation
Xu, T., Liu, Q., and Peng, J · 2017
Later among the works it cites.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Fang, C., Li, C. J., Lin, Z., and Zhang, T · 2018
Later among the works it cites.
A simple proximal stochastic gradient method for nonsmooth nonconvex optimization
Li, Z. and Li, J · 2018
Later among the works it cites.
Policy optimization via importance sampling
Metelli, A. M., Papini, M., Faccio, F., and Restelli, M · 2018
Later among the works it cites.
Stochastic variance-reduced policy gradient
Papini, M., Binaghi, D., Canonaco, G., Pirotta, M., and Restelli, M · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
The mirage of action-dependent baselines in reinforcement learning
Tucker, G., Bhupatiraju, S., Gu, S., Turner, R., Ghahramani, Z., and Levine, S · 2018
Later among the works it cites.