Fetching the paper…
Reading the bibliography…
Policy gradient is a generic and flexible reinforcement learning approach that generally enjoys simplicity in analysis, implementation, and deployment.
Stochastic optimization
V. M. Aleksandrov, V. I. Sysoyev, and V. V. Shemeneva · 1968
Earlier work this paper cites.
Some problems in monte carlo optimization
R. Y. Rubinstein · 1969
Earlier work this paper cites.
The complexity of markov decision processes
Christos H Papadimitriou and John N Tsitsiklis · 1987
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Advantage updating
Leemon C Baird III · 1993
Earlier work this paper cites.
Memoryless policies: Theoretical limitations and practical results
Michael L Littman · 1994
Earlier work this paper cites.
Learning without state-estimation in partially observable markovian decision processes
Satinder P Singh, Tommi Jaakkola, and Michael I Jordan · 1994
Earlier work this paper cites.
On the undecidability of probabilistic planning and infinite-horizon partially observable markov decision problems
Omid Madani, Steve Hanks, and Anne Condon · 1999
Earlier work this paper cites.
Experimental results on learning stochastic memoryless policies for partially observable markov decision processes
John K Williams and Satinder P Singh · 1999
Earlier work this paper cites.
Reinforcement learning in pomdp’s via direct gradient ascent
Jonathan Baxter and Peter L. Bartlett · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Cited alongside, same era.
Riemannian manifolds: an introduction to curvature , volume 176
John M Lee · 2006
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Cited alongside, same era.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
On the computational complexity of stochastic controller optimization in pomdps
Nikos Vlassis, Michael L Littman, and David Barber · 2012
Guido Montufar, Keyan Ghazi-Zahedi, and Nihat Ay · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Later among the works it cites.
Information geometry and its applications
Shun-ichi Amari · 2016
Later among the works it cites.
Reinforcement learning of pomdps using spectral methods
Kamyar Azizzadenesheli, Alessandro Lazaric, and Animashree Anandkumar · 2016
Later among the works it cites.
Combating reinforcement learning’s sisyphean curse with intrinsic fear
Zachary C Lipton, Kamyar Azizzadenesheli, Abhishek Kumar, Lihong Li, Jianfeng Gao, and Li Deng · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Bias in natural actor-critic algorithms
Philip Thomas · 2014
Cited alongside, same era.
The dependence of effective planning horizon on model accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Later among the works it cites.
Experimental results: Reinforcement learning of pomdps using spectral methods
Kamyar Azizzadenesheli, Alessandro Lazaric, and Animashree Anandkumar · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi · 2018
Closest in time.
Politex: Regret bounds for policy iteration using expert prediction
Yasin Abbasi-Yadkori, Peter Bartlett, Kush Bhatia, Nevena Lazic, Csaba Szepesvari, and Gellért Weisz · 2019
Closest in time.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2019
Closest in time.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Closest in time.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Closest in time.