Fetching the paper…
Reading the bibliography…
We propose expected policy gradients (EPG), which unify stochastic policy gradients (SPG) and deterministic policy gradients (DPG) for reinforcement learning.
On the Theory of the Brownian Motion
George E Uhlenbeck and Leonard S Ornstein · 1930
Earlier work this paper cites.
The Generalized Weierstrass Approximation Theorem
Marshall H Stone · 1948
Earlier work this paper cites.
Reinforcement Learning Applied to Linear Quadratic Regulation
Steven J. Bradtke · 1993
Earlier work this paper cites.
On-Line Q-Learning Using Connectionist Systems
Gavin A Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
Residual Algorithms: Reinforcement Learning with Function Approximation
Leemon Baird et al · 1995
Earlier work this paper cites.
Generalization in Reinforcement Learning: Successful Examples Using Sparse Coarse Coding
Richard S Sutton · 1996
Earlier work this paper cites.
Natural Gradient Works Efficiently in Learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Linear Quadratic Regulation using Reinforcement Learning
ten Hagen, S.H.G., Kröse, B.J.A., van der Broek, W., Verdenius, F., and Amsterdam Machine Learning lab (IVI, FNWI) · 1998
Earlier work this paper cites.
A Natural Policy Gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Least-Squares Policy Iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Policy Gradient Methods for Robot Control
Jan Peters, Sethu Vijaykumar, and Stefan Schaal · 2003
Earlier work this paper cites.
Iterative Linear Quadratic Regulator Design for Nonlinear Biological Movement Systems
Weiwei Li and Emanuel Todorov · 2004
Earlier work this paper cites.
Mathematical Statistics, Basic Ideas and Selected Topics, Vol. 1, (2nd Edition)
Peter Bickel and Kjell Doksum · 2006
Earlier work this paper cites.
Policy Gradient Methods for Robotics
Jan Peters and Stefan Schaal · 2006
Earlier work this paper cites.
Incremental Natural Actor-Critic Algorithms
Shalabh Bhatnagar, Mohammad Ghavamzadeh, Mark Lee, and Richard S Sutton · 2008
Earlier work this paper cites.
A Theoretical and Empirical Analysis of Expected Sarsa
Harm van Seijen, Hado van Hasselt, Shimon Whiteson, and Marco Wiering · 2009
Earlier work this paper cites.
Relative Entropy Policy Search
Jan Peters, Katharina Mülling, and Yasemin Altun · 2010
Earlier work this paper cites.
Optimal Estimation of Dynamic Systems
John L Crassidis and John L Junkins · 2011
Earlier work this paper cites.
Thomas Degris, Martha White, and Richard S Sutton · 2012
Earlier work this paper cites.
Value-Gradient Learning
Michael Fairbank and Eduardo Alonso · 2012
Cited alongside, same era.
A Unifying Perspective of Parametric Policy Search Methods for Markov Decision Processes
Thomas Furmston and David Barber · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
MuJoCo: A Physics Engine for Model-Based Control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Adaptive Step-Size for Policy Gradient Methods
Matteo Pirotta, Marcello Restelli, and Luca Bascetta · 2013
Cited alongside, same era.
Value-Gradient Learning
Michael Fairbank · 2014
Cited alongside, same era.
Multi-Objective Reinforcement Learning through Continuous Pareto Manifold Approximation
Simone Parisi, Matteo Pirotta, and Marcello Restelli · 2016
Later among the works it cites.
Nonlinear Kalman Filters Explained: A Tutorial on Moment Computations and Sigma Point Methods
Michael Roth, Gustaf Hendeby, and Fredrik Gustafsson · 2016
Later among the works it cites.
Mean Actor Critic
K. Asadi, C. Allen, M. Roderick, A.-r. Mohamed, G. Konidaris, and M. Littman · 2017
Later among the works it cites.
OpenAI Baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Later among the works it cites.
Interpolated Policy Gradient: Merging on-Policy and off-Policy Gradient Estimation for Deep Reinforcement Learning
Shixiang Shane Gu, Timothy Lillicrap, Richard E. Turner, Zoubin Ghahramani, Bernhard Schölkopf, and Sergey Levine · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 2014
Cited alongside, same era.
Deterministic Policy Gradient Algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015, 2015
Martın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, and Matthieu Devin · 2015
Cited alongside, same era.
Model-based relative entropy stochastic search
Abbas Abdolmaleki, Rudolf Lioutikov, Jan R. Peters, Nuno Lau, Luis Pualo Reis, and Gerhard Neumann · 2015
Cited alongside, same era.
Learning Continuous Control Policies by Stochastic Value Gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Riashat Islam, Peter Henderson, Maziar Gomrokchi, and Doina Precup · 2017
Later among the works it cites.
Bridging the Gap Between Value and Policy Based Reinforcement Learning
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2017
Later among the works it cites.
A Unified View of Entropy-Regularized Markov Decision Processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Later among the works it cites.
Combining policy gradient and q-learning
Brendan O’Donoghue, Rémi Munos, Koray Kavukcuoglu, and Volodymyr Mnih · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Philip S. Thomas and Emma Brunskill · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Rémi Munos, Nicolas Heess, and Martin A. Riedmiller · 2018
Closest in time.
Expected Policy Gradients
Kamil Ciosek and Shimon Whiteson · 2018
Closest in time.
Addressing Function Approximation Error in Actor-Critic Methods
Scott Fujimoto, Herke van Hoof, and Dave Meger · 2018
Closest in time.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Closest in time.
Trust-pcl: An off-policy trust region method for continuous control
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2018
Closest in time.
Total stochastic gradient algorithms and applications in reinforcement learning
Paavo Parmas · 2018
Closest in time.
Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines
Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M. Bayen, Sham Kakade, Igor Mordatch, and Pieter Abbeel · 2018
Closest in time.