Understand
Guided policy search algorithms can be used to optimize complex nonlinear policies, such as deep neural networks, without directly computing policy gradients in the high-dimensional parameter space.
- Instead, these methods use supervised learning to train the policy to mimic a "teacher" algorithm, such as a trajectory optimizer or a trajectory-centric reinforcement learning method.
- Guided policy search methods provide asymptotic local convergence guarantees by construction, but it is not clear how much the policy improves within a small, finite number of iterations.
- We show that guided policy search algorithms can be interpreted as an approximate variant of mirror descent, where the projection onto the constraint manifold is not exact.
Built on
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. Williams · 1992
Earlier work this paper cites.
Covariant policy search
J. A. Bagnell and J. Schneider · 2003
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
A. Beck and M. Teboulle · 2003
Earlier work this paper cites.
Iterative linear quadratic regulator design for nonlinear biological movement systems
W. Li and E. Todorov · 2004
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Relative entropy policy search
J. Peters, K. Mülling, and Y. Altün · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and A. Bagnell · 2011
Earlier work this paper cites.
Similar
A survey on policy search for robotics
M. Deisenroth, G. Neumann, and J. Peters · 2013
Cited alongside, same era.
Variational policy search via trajectory optimization
S. Levine and V. Koltun · 2013
Cited alongside, same era.
Learning monocular reactive UAV control in cluttered natural environments
S. Ross, N. Melik-Barkhudarov, K. Shaurya Shankar, A. Wendel, D. Dey, J. A. Bagnell, and M. Hebert · 2013
Cited alongside, same era.
Deep learning for real-time Atari game play using offline Monte-Carlo tree search planning
X. Guo, S. Singh, H. Lee, R. L. Lewis, and X. Wang · 2014
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
S. Levine and P. Abbeel · 2014
Cited alongside, same era.
Combining the benefits of function approximation and trajectory optimization
I. Mordatch and E. Todorov · 2014
Cited alongside, same era.
Then
Learning contact-rich manipulation skills with guided policy search
S. Levine, N. Wagener, and P. Abbeel · 2015
Later among the works it cites.
Interactive control of diverse complex characters with neural networks
I. Mordatch, K. Lowrey, G. Andrew, Z. Popovic, and E. Todorov · 2015
Later among the works it cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Moritz, M. Jordan, and P. Abbeel · 2015
Later among the works it cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Closest in time.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Closest in time.
Learning deep control policies for autonomous aerial vehicles with mpc-guided policy search
T. Zhang, G. Kahn, S. Levine, and P. Abbeel · 2016
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…