Fetching the paper…
Reading the bibliography…
We adapt recent tools developed for the analysis of Stochastic Gradient Descent (SGD) in non-convex optimization to obtain convergence and sample complexity guarantees for the vanilla policy gradient (PG).
Sur le problème de la division
Stanisław Łojasiewicz · 1959
Earlier work this paper cites.
Une propriété topologique des sous-ensembles analytiques réels
Stanisław Łojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
Boris T. Polyak · 1963
Earlier work this paper cites.
Pseudogradient adaptation and training algorithms
Boris Polyak and Y.Z. Tsypkin · 1973
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J. Williams and Jing Peng · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
On gradients of functions definable in o-minimal structures
Krzysztof Kurdyka · 1998
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
J. Baxter and P. L. Bartlett · 2001
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Sensitivity and convergence of uniformly ergodic markov chains
A. Yu. Mitrophanov · 2005
Earlier work this paper cites.
Natural actor–critic algorithms
Shalabh Bhatnagar, Richard S. Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2009
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition, 2013
Mark Schmidt and Nicolas Le Roux · 2013
Earlier work this paper cites.
Policy gradient in lipschitz markov decision processes
Matteo Pirotta, Marcello Restelli, and Luca Bascetta · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
First-Order Methods in Optimization
Amir Beck · 2017
Earlier work this paper cites.
Bridging the gap between value and policy based reinforcement learning
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2017
Earlier work this paper cites.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Lam M. Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Cited alongside, same era.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
A hybrid stochastic policy gradient algorithm for reinforcement learning
Nhan Pham, Lam Nguyen, Dzung Phan, Phuong Ha Nguyen, Marten van Dijk, and Quoc Tran-Dinh · 2020
Later among the works it cites.
The impact of neural network overparameterization on gradient confusion and stochastic gradient descent
Karthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang, and Tom Goldstein · 2020
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Later among the works it cites.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2020
Later among the works it cites.
Stochastic recursive momentum for policy gradient methods, 2020
Huizhuo Yuan, Xiangru Lian, Ji Liu, and Yuren Zhou · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Stochastic variance-reduced policy gradient
Matteo Papini, Damiano Binaghi, Giuseppe Canonaco, Matteo Pirotta, and Marcello Restelli · 2018
Cited alongside, same era.
Global optimality guarantees for policy gradient methods, 2019
Jalaj Bhandari and Daniel Russo · 2019
Cited alongside, same era.
Momentum-based variance reduction in non-convex SGD
Ashok Cutkosky and Francesco Orabona · 2019
Cited alongside, same era.
SGD: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Cited alongside, same era.
Stochastic gradient descent for nonconvex learning without bounded gradient assumptions
Yunwen Lei, Ting Hu, Guiying Li, and Ke Tang · 2019
Cited alongside, same era.
On principled entropy exploration in policy optimization
Jincheng Mei, Chenjun Xiao, Ruitong Huang, Dale Schuurmans, and Martin Müller · 2019
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M. Kakade, Jason D. Lee, and Gaurav Mahajan · 2021
Closest in time.
Beyond exact gradients: Convergence of stochastic soft-max policy gradient methods with entropy regularization, 2021
Yuhao Ding, Junzi Zhang, and Javad Lavaei · 2021
Closest in time.
Stochastic quasi-gradient methods: variance reduction via Jacobian sketching
Robert M. Gower, Peter Richtárik, and Francis Bach · 2021
Closest in time.
On the sample complexity of actor-critic method for reinforcement learning with function approximation, 2021
Harshat Kumar, Alec Koppel, and Alejandro Ribeiro · 2021
Closest in time.
Softmax policy gradient methods can take exponential time to converge
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2021
Closest in time.
Leveraging non-uniformity in first-order non-convex optimization
Jincheng Mei, Yue Gao, Bo Dai, Csaba Szepesvari, and Dale Schuurmans · 2021
Closest in time.
A hybrid stochastic optimization framework for composite nonconvex optimization
Quoc Tran-Dinh, Nhan H. Pham, Dzung T. Phan, and Lam M. Nguyen · 2021
Closest in time.
Non-asymptotic convergence of adam-type reinforcement learning algorithms under markovian sampling
Huaqing Xiong, Tengyu Xu, Yingbin Liang, and Wei Zhang · 2021
Closest in time.
On the global optimum convergence of momentum-based policy gradient
Yuhao Ding, Junzi Zhang, and Javad Lavaei · 2022
Closest in time.
Bregman gradient policy optimization
Feihu Huang, Shangqian Gao, and Heng Huang · 2022
Closest in time.
Smoothing policies and safe policy gradients
Matteo Papini, Matteo Pirotta, and Marcello Restelli · 2022
Closest in time.
Mirror descent policy optimization
Manan Tomar, Lior Shani, Yonathan Efroni, and Mohammad Ghavamzadeh · 2022
Closest in time.
A general class of surrogate functions for stable and efficient reinforcement learning
Sharan Vaswani, Olivier Bachem, Simone Totaro, Robert Müller, Shivam Garg, Matthieu Geist, Marlos C. Machado, Pablo Samuel Castro, and Nicolas Le Roux · 2022
Closest in time.
On the convergence rates of policy gradient methods
Lin Xiao · 2022
Closest in time.
Policy optimization with stochastic mirror descent
Long Yang, Yu Zhang, Gang Zheng, Qian Zheng, Pengfei Li, Jianhang Huang, and Gang Pan · 2022
Closest in time.