Policy gradient in Lipschitz Markov Decision Processes
Matteo Pirotta, Marcello Restelli, and Luca Bascetta · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Later among the works it cites.
Openai gym
Original
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Original
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
Provably efficient maximum entropy exploration
Original
Elad Hazan, Sham M Kakade, Karan Singh, and Abby Van Soest · 2018
Later among the works it cites.
Stochastic variance-reduced policy gradient
Original
Matteo Papini, Damiano Binaghi, Giuseppe Canonaco, Matteo Pirotta, and Marcello Restelli · 2018
Later among the works it cites.
Optimality and approximation with policy gradient methods in Markov decision processes
Original
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2019
Later among the works it cites.
Global optimality guarantees for policy gradient methods
Original
Jalaj Bhandari and Daniel Russo · 2019
Later among the works it cites.
LQR through the lens of first order methods: Discrete-time case
Original
Jingjing Bu, Afshin Mesbahi, Maryam Fazel, and Mehran Mesbahi · 2019
Later among the works it cites.
Efficiency of minimizing compositions of convex functions and smooth maps
Dmitriy Drusvyatskiy and Courtney Paquette · 2019
Later among the works it cites.
Neural proximal/trust region policy optimization attains globally optimal policy
Original
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Later among the works it cites.
Neural policy gradient methods: Global optimality and rates of convergence
Original
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Later among the works it cites.
Sample efficient policy gradient methods with recursive variance reduction
Original
Pan Xu, Felicia Gao, and Quanquan Gu · 2019
Later among the works it cites.
Global convergence of policy gradient methods to (almost) locally optimal policies
Original
Kaiqing Zhang, Alec Koppel, Hao Zhu, and Tamer Başar · 2019
Later among the works it cites.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2020
Closest in time.
On the global convergence rates of softmax policy gradient methods
Original
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Closest in time.