Fetching the paper…
Reading the bibliography…
Policy optimization is a fundamental principle for designing reinforcement learning algorithms, and one example is the proximal policy optimization algorithm with a clipped surrogate objective (PPO-Clip), which has been popularly used in deep reinforcement learning due to its simplicity and effectiveness.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Large margin classification using the perceptron algorithm
Yoav Freund and Robert E. Schapire · 1999
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
Satinder Singh, Tommi Jaakkola, Michael L Littman, and Csaba Szepesvári · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham M. Kakade and John Langford · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Reinforcement learning as classification: Leveraging modern classifiers
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Fitted q-iteration in continuous action-space mdps
András Antos, Csaba Szepesvári, and Rémi Munos · 2007
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Earlier work this paper cites.
Analysis of a classification-based policy iteration algorithm
Alessandro Lazaric, Mohammad Ghavamzadeh, and Remi Munos · 2010
Earlier work this paper cites.
Classification-based approximate policy iteration: Experiments and extended discussions
Amir-massoud Farahmand, Doina Precup, André Barreto, and Mohammad Ghavamzadeh · 2014
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Regularized policy iteration with nonparametric function spaces
Amir-massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Mastering complex control in moba games with deep reinforcement learning
Deheng Ye, Zhao Liu, Mingfei Sun, Bei Shi, Peilin Zhao, Hao Wu, Hongsheng Yu, Shaojie Yang, Xipeng Wu, Qingwei Guo, et al · 2020
Later among the works it cites.
Proximal policy gradient: Ppo with policy gradient
Ju-Seung Byun, Byungmoon Kim, and Huamin Wang · 2020
Later among the works it cites.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Later among the works it cites.
On the global convergence rates of softmax policy gradient methods
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Later among the works it cites.
An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods
Yanli Liu, Kaiqing Zhang, Tamer Basar, and Wotao Yin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Boosted fitted q-iteration
Samuele Tosatto, Matteo Pirotta, Carlo d’Eramo, and Marcello Restelli · 2017
Cited alongside, same era.
An adaptive clipping approach for proximal policy optimization
Gang Chen, Yiming Peng, and Mengjie Zhang · 2018
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2019
Cited alongside, same era.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Cited alongside, same era.
Neural trust region/proximal policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Cited alongside, same era.
MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments
Kenny Young and Tian Tian · 2019
Cited alongside, same era.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Later among the works it cites.
Rethinking deep policy gradients via state-wise policy improvement
Kai-Chun Hu, Ping-Chun Hsieh, Ting Han Wei, and I-Chen Wu · 2020
Later among the works it cites.
Low-level autonomous control and tracking of quadrotor using reinforcement learning
Chen-Huan Pi, Kai-Chun Hu, Stone Cheng, and I-Chen Wu · 2020
Later among the works it cites.
RL Baselines 3 Zoo
Antonin Raffin · 2020
Later among the works it cites.
Global convergence of policy gradient for linear-quadratic mean-field control/game in continuous time
Weichen Wang, Jiequn Han, Zhuoran Yang, and Zhaoran Wang · 2021
Closest in time.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Johan Samir Obando Ceron and Pablo Samuel Castro · 2021
Closest in time.