Fetching the paper…
Reading the bibliography…
Policy gradient methods have demonstrated success in reinforcement learning tasks that have high-dimensional continuous state and action spaces.
A Markovian decision process
R. Bellman · 1957
Earlier work this paper cites.
Statistical physics (course of theoretical physics vol 5)
L. Landau and E. Lifshitz · 1958
Earlier work this paper cites.
A course in simulation
S. M. Ross · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
D. P. Bertsekas, D. P. Bertsekas, D. P. Bertsekas, and D. P. Bertsekas · 1995
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
S. P. Singh and R. S. Sutton · 1996
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
An analysis of actor-critic algorithms using eligibility traces: reinforcement learning with imperfect value functions
H. Kimura, S. Kobayashi, et al · 2000
Earlier work this paper cites.
A course in probability theory
K. L. Chung · 2001
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
J. Baxter and P. L. Bartlett · 2001
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
A. Beck and M. Teboulle · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade et al · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
E. Greensmith, P. L. Bartlett, and J. Baxter · 2004
Cited alongside, same era.
Pattern recognition and machine learning
C. M. Bishop · 2006
Cited alongside, same era.
Natural actor-critic
J. Peters and S. Schaal · 2008
Cited alongside, same era.
On a connection between importance sampling and the likelihood ratio policy gradient
T. Jie and P. Abbeel · 2010
Cited alongside, same era.
Monte Carlo theory, methods and examples
A. B. Owen · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2017
Later among the works it cites.
Fast policy learning through imitation and reinforcement
C.-A. Cheng, X. Yan, N. Wagener, and B. Boots · 2018
Later among the works it cites.
Truncated horizon policy search: Combining reinforcement learning & imitation learning
W. Sun, J. A. Bagnell, and B. Boots · 2018
Later among the works it cites.
Action-depedent control variates for policy optimization via stein’s identity
H. Liu, Y. Feng, Y. Mao, D. Zhou, J. Peng, and Q. Liu · 2018
Later among the works it cites.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
W. Grathwohl, D. Choi, Y. Wu, G. Roeder, and D. Duvenaud · 2018
Later among the works it cites.
The mirage of action-dependent baselines in reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Variance reduction for stochastic gradient optimization
C. Wang, X. Chen, A. J. Smola, and E. P. Xing · 2013
Cited alongside, same era.
Bias in natural actor-critic algorithms
P. Thomas · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
S. Ghadimi, G. Lan, and H. Zhang · 2016
Cited alongside, same era.
G. Tucker, S. Bhupatiraju, S. Gu, R. E. Turner, Z. Ghahramani, and S. Levine · 2018
Later among the works it cites.
Reward-estimation variance elimination in sequential decision processes
S. Pankov · 2018
Later among the works it cites.
Variance reduction for policy gradient with action-dependent factorized baselines
C. Wu, A. Rajeswaran, Y. Duan, V. Kumar, A. M. Bayen, S. Kakade, I. Mordatch, and P. Abbeel · 2018
Later among the works it cites.
Expected policy gradients for reinforcement learning
K. Ciosek and S. Whiteson · 2018
Later among the works it cites.
DART: Dynamic animation and robotics toolkit
J. Lee, M. X. Grey, S. Ha, T. Kunz, S. Jain, Y. Ye, S. S. Srinivasa, M. Stilman, and C. K. Liu · 2018
Later among the works it cites.
Predictor-corrector policy optimization
C.-A. Cheng, X. Yan, N. Ratliff, and B. Boots · 2019
Closest in time.
Policy optimization with stochastic mirror descent
L. Yang and Y. Zhang · 2019
Closest in time.
Beyond the one step greedy approach in reinforcement learning
Y. Efroni, G. Dalal, B. Scherrer, and S. Mannor · 2019
Closest in time.
Contrasting exploration in parameter and action space: A zeroth-order optimization perspective
A. Vemula, W. Sun, and J. A. Bagnell · 2019
Closest in time.