Fetching the paper…
Reading the bibliography…
Policy gradient methods are very attractive in reinforcement learning due to their model-free nature and convergence guarantees.
Stochastic optimization
VM Aleksandrov, VI Sysoyev, and VV Shemenev · 1968
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Richard Stuart Sutton · 1984
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Optimal control and estimation
Robert F Stengel · 1994
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
Policy gradient methods for robotics
Jan Peters and Stefan Schaal · 2006
Cited alongside, same era.
Controlled diffusion processes
Nikolaj Vladimirovič Krylov · 2008
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Will Grathwohl, Dami Choi, Yuhuai Wu, Geoff Roeder, and David Duvenaud · 2017
Later among the works it cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine · 2017
Later among the works it cites.
Action-dependent control variates for policy optimization via stein identity
Hao Liu, Yihao Feng, Yi Mao, Dengyong Zhou, Jian Peng, and Qiang Liu · 2018
Closest in time.
The mirage of action-dependent baselines in reinforcement learning
George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard E Turner, Zoubin Ghahramani, and Sergey Levine · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E Turner, and Sergey Levine · 2016
Cited alongside, same era.
Closest in time.
Variance reduction for policy gradient with action-dependent factorized baselines
Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M Bayen, Sham Kakade, Igor Mordatch, and Pieter Abbeel · 2018
Closest in time.