Fetching the paper…
Reading the bibliography…
We present a unified framework for learning continuous control policies using backpropagation.
Differential dynamic programming
D. H. Jacobson and D. Q. Mayne · 1970
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Identification and control of dynamical systems using neural networks
K. S. Narendra and K. Parthasarathy · 1990
Earlier work this paper cites.
Neural networks for self-learning control systems
D. H. Nguyen and B. Widrow · 1990
Earlier work this paper cites.
A menu of designs for reinforcement learning over time
P. J Werbos · 1990
Earlier work this paper cites.
Forward models: Supervised learning with a distal teacher
M. I. Jordan and D. E. Rumelhart · 1992
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L.-J. Lin · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Learning without state-estimation in partially observable Markovian decision processes
S. P. Singh · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
L. Baird · 1995
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R.S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Reinforcement learning using neural networks, with applications to motor control
R. Coulom · 2002
Cited alongside, same era.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Cited alongside, same era.
Using inaccurate models in reinforcement learning
P. Abbeel, M. Quigley, and A. Y. Ng · 2006
Cited alongside, same era.
Policy gradient in continuous time
R. Munos · 2006
Cited alongside, same era.
Receding horizon differential dynamic programming
Y. Tassa, T. Erez, and W.D. Smart · 2008
Cited alongside, same era.
A cat-like robot real-time learning to run
P. Wawrzyński · 2009
Cited alongside, same era.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Later among the works it cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Later among the works it cites.
Value-gradient learning
M. Fairbank · 2014
Later among the works it cites.
Learning neural network policies with guided policy search under unknown dynamics
S. Levine and P. Abbeel · 2014
Later among the works it cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Later among the works it cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Real-time reinforcement learning by sequential actor–critics and experience replay
P. Wawrzyński · 2009
Cited alongside, same era.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Cited alongside, same era.
Efficient robust policy optimization
C. G. Atkeson · 2012
Cited alongside, same era.
Value-gradient learning
M. Fairbank and E. Alonso · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Later among the works it cites.
Compatible value gradients for reinforcement learning of continuous deep policies
D. Balduzzi and M. Ghifary · 2015
Closest in time.
Online Model Learning Algorithms for Actor-Critic Control
I. Grondman · 2015
Closest in time.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Closest in time.
Trust region policy optimization
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel · 2015
Closest in time.