Fetching the paper…
Reading the bibliography…
Recently, the impressive empirical success of policy gradient (PG) methods has catalyzed the development of their theoretical foundations.
Une propriété topologique des sous-ensembles analytiques réels
Stanislaw Lojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for minimizing functionals
Boris Teodorovich Polyak · 1963
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Information theory and statistics
Solomon Kullback · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J.N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
On gradients of functions definable in o-minimal structures
Krzysztof Kurdyka · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
An analysis of reinforcement learning with function approximation
Francisco S Melo, Sean P Meyn, and M Isabel Ribeiro · 2008
Earlier work this paper cites.
On the convergence of the proximal algorithm for nonsmooth functions involving analytic features
Hedy Attouch and Jérôme Bolte · 2009
Earlier work this paper cites.
The exponential family and statistical applications
Anirban DasGupta and Anirban DasGupta · 2011
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Beyond Convexity: Stochastic Quasi-Convex Optimization
Elad Hazan, Kfir Y. Levy, and Shai Shalev-Shwartz · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Earlier work this paper cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Earlier work this paper cites.
Stochastic Heavy ball
Sébastien Gadat, Fabien Panloup, and Sofiane Saadane · 2018
Earlier work this paper cites.
Stochastic variance-reduced policy gradient
Matteo Papini, Damiano Binaghi, Giuseppe Canonaco, Matteo Pirotta, and Marcello Restelli · 2018
Cited alongside, same era.
Reducing the variance in online optimization by transporting past gradients
Sébastien Arnold, Pierre-Antoine Manzagol, Reza Babanezhad Harikandeh, Ioannis Mitliagkas, and Nicolas Le Roux · 2019
Cited alongside, same era.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Cited alongside, same era.
Momentum-based variance reduction in non-convex sgd
Ashok Cutkosky and Francesco Orabona · 2019
Cited alongside, same era.
Garage: A toolkit for reproducible reinforcement learning research
The garage contributors · 2019
Cited alongside, same era.
Hessian aided policy gradient
Zebang Shen, Alejandro Ribeiro, Hamed Hassani, Hui Qian, and Chao Mi · 2019
Yuhao Ding, Junzi Zhang, and Javad Lavaei · 2021
Later among the works it cites.
Optimizing Static Linear Feedback: Gradient Method
Ilyas Fatkhullin and Boris Polyak · 2021
Later among the works it cites.
Convergence rates and approximation results for sgd and its continuous-time counterpart
Xavier Fontaine, Valentin De Bortoli, and Alain Durmus · 2021
Later among the works it cites.
Page: A simple and optimal probabilistic gradient estimator for nonconvex optimization
Zhize Li, Hongyan Bao, Xiangliang Zhang, and Peter Richtarik · 2021
Later among the works it cites.
On finite-time convergence of actor-critic algorithm
Shuang Qiu, Zhuoran Yang, Jieping Ye, and Zhaoran Wang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unified optimal analysis of the (stochastic) gradient method
Sebastian U. Stich · 2019
Cited alongside, same era.
Second-order information in non-convex stochastic optimization: Power and limitations
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Ayush Sekhari, and Karthik Sridharan · 2020
Cited alongside, same era.
Momentum improves normalized SGD
Ashok Cutkosky and Harsh Mehta · 2020
Cited alongside, same era.
Mingyi Hong, Hoi-To Wai, Zhaoran Wang, and Zhuoran Yang · 2020
Cited alongside, same era.
Momentum-based policy gradient methods
Feihu Huang, Shangqian Gao, Jian Pei, and Heng Huang · 2020
Cited alongside, same era.
An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods
Yanli Liu, Kaiqing Zhang, Tamer Basar, and Wotao Yin · 2020
Cited alongside, same era.
Hoang Tran and Ashok Cutkosky · 2021
Later among the works it cites.
Sample complexity of policy gradient finding second-order stationary points
Long Yang, Qian Zheng, and Gang Pan · 2021
Later among the works it cites.
Linear convergence for natural policy gradient with log-linear policy parametrization
Carlo Alfano and Patrick Rebeschini · 2022
Later among the works it cites.
On the hidden biases of policy mirror ascent in continuous action spaces
Amrit Singh Bedi, Souradip Chakraborty, Anjaly Parayil, Brian M Sadler, Pratap Tokekar, and Alec Koppel · 2022
Later among the works it cites.
Sample complexity of policy-based methods under off-policy sampling and linear function approximation
Zaiwei Chen and Siva Theja Maguluri · 2022
Later among the works it cites.
Finite-sample analysis of off-policy natural actor–critic with linear function approximation
Zaiwei Chen, Sajad Khodadadian, and Siva Theja Maguluri · 2022
Later among the works it cites.
On the global optimum convergence of momentum-based policy gradient
Yuhao Ding, Junzi Zhang, and Javad Lavaei · 2022
Later among the works it cites.
Sharp analysis of stochastic optimization under global Kurdyka-łojasiewicz inequality
Ilyas Fatkhullin, Jalal Etesami, Niao He, and Negar Kiyavash · 2022
Later among the works it cites.
PAGE-PG: A simple and loopless variance-reduced policy gradient method with probabilistic gradient estimation
Matilde Gargiani, Andrea Zanelli, Andrea Martinelli, Tyler Summers, and John Lygeros · 2022
Later among the works it cites.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Guanghui Lan · 2022
Later among the works it cites.
Stochastic second-order methods provably beat sgd for gradient-dominated functions
Saeed Masiha, Saber Salehkaleybar, Niao He, Negar Kiyavash, and Patrick Thiran · 2022
Later among the works it cites.
Adaptive momentum-based policy gradient with second-order information
Saber Salehkaleybar, Sadegh Khorasani, Negar Kiyavash, Niao He, and Patrick Thiran · 2022
Later among the works it cites.
Convergence Rates of Non-Convex Stochastic Gradient Descent Under a Generic Lojasiewicz Condition and Local Smoothness
Kevin Scaman, Cedric Malherbe, and Ludovic Dos Santos · 2022
Later among the works it cites.
On the convergence rates of policy gradient methods
Lin Xiao · 2022
Later among the works it cites.
Convergence and optimality of policy gradient methods in weakly smooth settings
Matthew S Zhang, Murat A Erdogdu, and Animesh Garg · 2022
Later among the works it cites.