Fetching the paper…
Reading the bibliography…
Policy gradient (PG) is widely used in reinforcement learning due to its scalability and good performance.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Earlier work this paper cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Adaptive step-size for policy gradient methods
Matteo Pirotta, Marcello Restelli, and Luca Bascetta · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Gradient descent efficiently finds the cubic-regularized non-convex newton step
Yair Carmon and John C Duchi · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Cited alongside, same era.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
A hybrid stochastic policy gradient algorithm for reinforcement learning
Nhan Pham, Lam Nguyen, Dzung Phan, Phuong Ha Nguyen, Marten Dijk, and Quoc Tran-Dinh · 2020
Later among the works it cites.
An improved convergence analysis of stochastic variance-reduced policy gradient
Pan Xu, Felicia Gao, and Quanquan Gu · 2020
Later among the works it cites.
Stochastic recursive variance-reduced cubic regularization methods
Dongruo Zhou and Quanquan Gu · 2020
Later among the works it cites.
On the global convergence of momentum-based policy gradient
Yuhao Ding, Junzi Zhang, and Javad Lavaei · 2021
Later among the works it cites.
Bregman gradient policy optimization
Feihu Huang, Shangqian Gao, and Heng Huang · 2021
Later among the works it cites.
Sample complexity of policy gradient finding second-order stationary points
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Stochastic variance-reduced policy gradient
Matteo Papini, Damiano Binaghi, Giuseppe Canonaco, Matteo Pirotta, and Marcello Restelli · 2018
Cited alongside, same era.
Stochastic cubic regularization for fast nonconvex optimization
Nilesh Tripuraneni, Mitchell Stern, Chi Jin, Jeffrey Regier, and Michael I Jordan · 2018
Cited alongside, same era.
Garage: A toolkit for reproducible reinforcement learning research
The garage contributors · 2019
Cited alongside, same era.
Hessian aided policy gradient
Zebang Shen, Alejandro Ribeiro, Hamed Hassani, Hui Qian, and Chao Mi · 2019
Cited alongside, same era.
Sample efficient policy gradient methods with recursive variance reduction
Pan Xu, Felicia Gao, and Quanquan Gu · 2019
Cited alongside, same era.
Momentum-based policy gradient methods
Feihu Huang, Shangqian Gao, Jian Pei, and Heng Huang · 2020
Cited alongside, same era.
Long Yang, Qian Zheng, and Gang Pan · 2021
Later among the works it cites.
On the convergence and sample efficiency of variance-reduced policy gradient method
Junyu Zhang, Chengzhuo Ni, Zheng Yu, Csaba Szepesvari, and Mengdi Wang · 2021
Later among the works it cites.
Matilde Gargiani, Andrea Zanelli, Andrea Martinelli, Tyler Summers, and John Lygeros · 2022
Later among the works it cites.
Stochastic second-order methods provably beat sgd for gradient-dominated functions
Saeed Masiha, Saber Salehkaleybar, Niao He, Negar Kiyavash, and Patrick Thiran · 2022
Later among the works it cites.
Stochastic cubic-regularized policy gradient method
Pengfei Wang, Hongyu Wang, and Nenggan Zheng · 2022
Later among the works it cites.
Policy optimization with stochastic mirror descent
Long Yang, Yu Zhang, Gang Zheng, Qian Zheng, Pengfei Li, Jianhang Huang, and Gang Pan · 2022
Later among the works it cites.
A cubic-regularized policy newton algorithm for reinforcement learning
Mizhaan P Maniyar, LA Prashanth, Akash Mondal, and Shalabh Bhatnagar · 2024
Closest in time.