Fetching the paper…
Reading the bibliography…
Variance-reduced gradient estimators for policy gradient methods have been one of the main focus of research in the reinforcement learning in recent years as they allow acceleration of the estimation process.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Earlier work this paper cites.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Spider: Near-optimal non-convex optimization via stochastic path integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Stochastic variance-reduced policy gradient
Matteo Papini, Damiano Binaghi, Giuseppe Canonaco, Matteo Pirotta, and Marcello Restelli · 2018
Cited alongside, same era.
Momentum-based variance reduction in non-convex sgd
Ashok Cutkosky and Francesco Orabona · 2019
Cited alongside, same era.
Garage: A toolkit for reproducible reinforcement learning research
The garage contributors · 2019
Cited alongside, same era.
Hessian aided policy gradient
Zebang Shen, Alejandro Ribeiro, Hamed Hassani, Hui Qian, and Chao Mi · 2019
Cited alongside, same era.
Spiderboost and momentum: Faster variance reduction algorithms
Zhe Wang, Kaiyi Ji, Yi Zhou, Yingbin Liang, and Vahid Tarokh · 2019
Cited alongside, same era.
Stochastic recursive momentum for policy gradient methods
Huizhuo Yuan, Xiangru Lian, Ji Liu, and Yuren Zhou · 2020
Later among the works it cites.
One sample stochastic frank-wolfe
Mingrui Zhang, Zebang Shen, Aryan Mokhtari, Hamed Hassani, and Amin Karbasi · 2020
Later among the works it cites.
Beyond variance reduction: Understanding the true impact of baselines on policy optimization
Wesley Chung, Valentin Thomas, Marlos C Machado, and Nicolas Le Roux · 2021
Later among the works it cites.
Storm+: Fully adaptive sgd with momentum for nonconvex optimization
Kfir Levy, Ali Kavis, and Volkan Cevher · 2021
Later among the works it cites.
Page: A simple and optimal probabilistic gradient estimator for nonconvex optimization
Zhize Li, Hongyan Bao, Xiangliang Zhang, and Peter Richtárik · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sample efficient policy gradient methods with recursive variance reduction
Pan Xu, Felicia Gao, and Quanquan Gu · 2019
Cited alongside, same era.
Policy optimization with stochastic mirror descent
Long Yang, Gang Zheng, Haotian Zhang, Yu Zhang, Qian Zheng, Jun Wen, and Gang Pan · 2019
Cited alongside, same era.
Second-order information in non-convex stochastic optimization: Power and limitations
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Ayush Sekhari, and Karthik Sridharan · 2020
Cited alongside, same era.
Momentum-based policy gradient methods
Feihu Huang, Shangqian Gao, Jian Pei, and Heng Huang · 2020
Cited alongside, same era.
An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods
Yanli Liu, Kaiqing Zhang, Tamer Basar, and Wotao Yin · 2020
Cited alongside, same era.
A hybrid stochastic policy gradient algorithm for reinforcement learning
Nhan Pham, Lam Nguyen, Dzung Phan, Phuong Ha Nguyen, Marten Dijk, and Quoc Tran-Dinh · 2020
Cited alongside, same era.
Junyu Zhang, Chengzhuo Ni, Zheng Yu, Csaba Szepesvari, and Mengdi Wang · 2021
Later among the works it cites.
Variance reduction for non-convex stochastic optimization: General analysis and new applications
Liang Zhang · 2021
Later among the works it cites.
On the global optimum convergence of momentum-based policy gradient
Yuhao Ding, Junzi Zhang, and Javad Lavaei · 2022
Closest in time.
Page-pg: A simple and loopless variance-reduced policy gradient method with probabilistic gradient estimation
Matilde Gargiani, Andrea Zanelli, Andrea Martinelli, Tyler Summers, and John Lygeros · 2022
Closest in time.
Bregman gradient policy optimization
Feihu Huang, Shangqian Gao, and Heng Huang · 2022
Closest in time.
Better sgd using second-order momentum
Hoang Tran and Ashok Cutkosky · 2022
Closest in time.
Policy optimization with stochastic mirror descent
Long Yang, Yu Zhang, Gang Zheng, Qian Zheng, Pengfei Li, Jianhang Huang, and Gang Pan · 2022
Closest in time.
Efficiently escaping saddle points for non-convex policy optimization
Sadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash, Niao He, and Matthias Grossglauser · 2023
Closest in time.