Fetching the paper…
Reading the bibliography…
Policy gradient methods are an appealing approach in reinforcement learning because they directly optimize the cumulative reward and can straightforwardly be used with nonlinear function approximators such as neural networks.
Principles of behavior
Hull, Clark · 1943
Earlier work this paper cites.
Steps toward artificial intelligence
Minsky, Marvin · 1961
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, Andrew G, Sutton, Richard S, and Anderson, Charles W · 1983
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
An analysis of actor/critic algorithms using eligibility traces: Reinforcement learning with imperfect value function
Kimura, Hajime and Kobayashi, Shigenobu · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, Andrew Y, Harada, Daishi, and Russell, Stuart · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, Richard S, McAllester, David A, Singh, Satinder P, and Mansour, Yishay · 1999
Earlier work this paper cites.
Numerical optimization
Wright, Stephen J and Nocedal, Jorge · 1999
Cited alongside, same era.
Reinforcement learning in POMDPs via direct gradient ascent
Baxter, Jonathan and Bartlett, Peter L · 2000
Cited alongside, same era.
On actor-critic algorithms
Konda, Vijay R and Tsitsiklis, John N · 2003
Cited alongside, same era.
Approximate gradient methods in policy-space optimization of markov reward processes
Marbach, Peter and Tsitsiklis, John N · 2003
Cited alongside, same era.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, Evan, Bartlett, Peter L, and Baxter, Jonathan · 2004
Cited alongside, same era.
Natural actor-critic
Peters, Jan and Schaal, Stefan · 2008
Cited alongside, same era.
Reinforcement learning in feedback control
Hafner, Roland and Riedmiller, Martin · 2011
Later among the works it cites.
Dynamic programming and optimal control , volume 2
Bertsekas, Dimitri P · 2012
Later among the works it cites.
Mujoco: A physics engine for model-based control
Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012
Later among the works it cites.
Bias in natural actor-critic algorithms
Thomas, Philip · 2014
Later among the works it cites.
Learning continuous control policies by stochastic value gradients
Heess, Nicolas, Wayne, Greg, Silver, David, Lillicrap, Timothy, Tassa, Yuval, and Erez, Tom · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convergent temporal-difference learning with arbitrary smooth function approximation
Bhatnagar, Shalabh, Precup, Doina, Silver, David, Sutton, Richard S, Maei, Hamid R, and Szepesvári, Csaba · 2009
Cited alongside, same era.
Real-time reinforcement learning by sequential actor–critics and experience replay
Wawrzyński, Paweł · 2009
Cited alongside, same era.
A natural policy gradient
Kakade, Sham
Cited in the paper.
Optimizing average reward using discounted rewards
Kakade, Sham
Cited in the paper.
Lillicrap, Timothy P, Hunt, Jonathan J, Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2015
Closest in time.
Trust region policy optimization
Schulman, John, Levine, Sergey, Moritz, Philipp, Jordan, Michael I, and Abbeel, Pieter · 2015
Closest in time.