Fetching the paper…
Reading the bibliography…
We present a policy search method for learning complex feedback control policies that map from high-dimensional sensory inputs to motor torques, for manipulation tasks with discontinuous contact dynamics.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shifts in position
K. Fukushima · 1980
Earlier work this paper cites.
Movement imitation with nonlinear dynamical systems in humanoid robots
J.A. Ijspeert, J. Nakanishi, and S. Schaal · 2002
Earlier work this paper cites.
Stochastic policy gradient reinforcement learning on a simple 3d biped
R. Tedrake, T.W. Zhang, and H.S. Seung · 2004
Earlier work this paper cites.
Policy gradient reinforcement learning for fast quadrupedal locomotion
N. Kohl and P. Stone · 2004
Earlier work this paper cites.
Learning perceptual coupling for motor primitives
J. Kober, B.J. Mohler, and J. Peters · 2008
Earlier work this paper cites.
Learning CPG-based biped locomotion with a policy gradient method: Application to a humanoid robot
G. Endo, J. Morimoto, T. Matsubara, J. Nakanishi, and G. Cheng · 2008
Earlier work this paper cites.
Learning and generalization of motor skills by learning from demonstration
P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal · 2009
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
E. Theodorou, J. Buchli, and S. Schaal · 2010
Earlier work this paper cites.
Reinforcement learning to adjust robot movements to new situations
J. Kober, E. Oztop, and J. Peters · 2010
Earlier work this paper cites.
Relative entropy policy search
J. Peters, K. Mülling, and Y. Altun · 2010
Earlier work this paper cites.
Learning to control a low-cost manipulator using data-efficient reinforcement learning
M.P. Deisenroth, C.E. Rasmussen, and D. Fox · 2011
Earlier work this paper cites.
Learning to grasp under uncertainty
F. Stulp, E. Theodorou, J. Buchli, and S. Schaal · 2011
Cited alongside, same era.
Learning force control policies for compliant manipulation
M. Kalakrishnan, L. Righetti, P. Pastor, and S. Schaal · 2011
Cited alongside, same era.
Path integral policy improvement with covariance matrix adaptation
F. Stulp and O. Sigaud · 2012
Cited alongside, same era.
A survey on policy search for robotics
M.P. Deisenroth, G. Neumann, and J. Peters · 2013
Cited alongside, same era.
Guided policy search
S. Levine and V. Koltun · 2013
Cited alongside, same era.
Learning robot tactile sensing for object manipulation
Y. Chebotar, O. Kroemer, and J. Peters · 2014
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Moritz, M. Jordan, and P. Abbeel · 2015
Later among the works it cites.
Learning contact-rich manipulation skills with guided policy search
S. Levine, N. Wagener, and P. Abbeel · 2015
Later among the works it cites.
Learning of non-parametric control policies with high-dimensional state features
H. van Hoof, J. Peters, and G. Neumann · 2015
Later among the works it cites.
Deep learning in neural networks: An overview
J. Schmidhuber · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Closest in time.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Levine and P. Abbeel · 2014
Cited alongside, same era.
Sample-based information-theoretic stochastic optimal control
R. Lioutikov, A. Paraschos, J. Peters, and G. Neumann · 2014
Cited alongside, same era.
Policy search for path integral control
V. Gómez, H.J. Kappen, J. Peters, and G. Neumann · 2014
Cited alongside, same era.
Evolving deep unsupervised convolutional networks for vision-based reinforcement learning
J. Koutník, J. Schmidhuber, and F.J. Gomez · 2014
Cited alongside, same era.
Joint training of a convolutional network and a graphical model for human pose estimation
J. Tompson, A. Jain, Y. LeCun, and C. Bregler · 2014
Cited alongside, same era.
Reinforcement learning with sequences of motion primitives for robust manipulation
F. Stulp, E. Theodorou, and S. Schaal
Cited in the paper.
Closest in time.
Guided policy search as approximate mirror descent
W. Montgomery and S. Levine · 2016
Closest in time.
Self-supervised regrasping using spatio-temporal tactile features and reinforcement learning
Y. Chebotar, K. Hausman, Z. Su, G.S. Sukhatme, and S. Schaal · 2016
Closest in time.
Combined optimization and reinforcement learning for manipulation skills
P. Englert and M. Toussaint · 2016
Closest in time.
Going further with point pair features
S. Hinterstoisser, V. Lepetit, N. Rajkumar, and K. Konolige · 2016
Closest in time.