Fetching the paper…
Reading the bibliography…
Autonomous learning of object manipulation skills can enable robots to acquire rich behavioral repertoires that scale to the variety of objects found in the real world.
V. Gullapalli, R. Grupen, and A. Barto, “Learning reactive admittance control,” in International Conference on Intelligent Robots and Systems (IROS) , 1992
1992
Earlier work this paper cites.
A. Ijspeert, J. Nakanishi, and S. Schaal, “Learning attractor landscapes for learning motor primitives,” in Advances in Neural Information Processing Systems (NIPS) , 2003
2003
Earlier work this paper cites.
J. A. Bagnell and J. Schneider, “Covariant policy search,” in International Joint Conference on Artificial Intelligence (IJCAI) , 2003
2003
Earlier work this paper cites.
N. Kohl and P. Stone, “Policy gradient reinforcement learning for fast quadrupedal locomotion,” in International Conference on Robotics and Automation (IROS) , 2004
2004
Earlier work this paper cites.
R. Tedrake, T. Zhang, and H. Seung, “Stochastic policy gradient reinforcement learning on a simple 3d biped,” in International Conference on Intelligent Robots and Systems (IROS) , 2004
2004
Earlier work this paper cites.
S. Boyd and L. Vandenberghe, Convex Optimization . New York, NY, USA: Cambridge University Press, 2004
2004
Earlier work this paper cites.
T. Geng, B. Porr, and F. Wörgötter, “Fast biped walking with a reflexive controller and realtime policy searching,” in Advances in Neural Information Processing Systems (NIPS) , 2006
2006
Earlier work this paper cites.
D. Bristow, M. Tharayil, and A. Alleyne, “A survey of iterative learning control,” Control Systems, IEEE , vol. 26, no. 3, 2006
2006
Earlier work this paper cites.
F. Guenter, M. Hersch, S. Calinon, and A. Billard, “Reinforcement learning for imitating constrained reaching movements,” Advanced Robotics , vol. 21, no. 13, pp. 1521–1544, 2007
2007
Earlier work this paper cites.
R. Hafner and M. Riedmiller, “Neural reinforcement learning controllers for a real robot application,” in International Conference on Robotics and Automation (ICRA) , 2007
2007
Earlier work this paper cites.
G. Endo, J. Morimoto, T. Matsubara, J. Nakanishi, and G. Cheng, “Learning CPG-based biped locomotion with a policy gradient method: Application to a humanoid robot,” International Journal of Robotic Research , vol. 27, no. 2, pp. 213–228, 2008
2008
Cited alongside, same era.
J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,” Neural Networks , vol. 21, no. 4, pp. 682–697, 2008
2008
Cited alongside, same era.
P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal, “Learning and generalization of motor skills by learning from demonstration,” in International Conference on Robotics and Automation (ICRA) , 2009
2009
Cited alongside, same era.
J. Kober, E. Oztop, and J. Peters, “Reinforcement learning to adjust robot movements to new situations,” in Robotics: Science and Systems , 2010
2010
Cited alongside, same era.
S. Ross, G. Gordon, and A. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” Journal of Machine Learning Research , vol. 15, pp. 627–635, 2011
2011
Later among the works it cites.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” International Journal of Robotic Research , vol. 32, no. 11, pp. 1238–1274, 2013
2013
Later among the works it cites.
M. Deisenroth, G. Neumann, and J. Peters, “A survey on policy search for robotics,” Foundations and Trends in Robotics , vol. 2, no. 1-2, pp. 1–142, 2013
2013
Later among the works it cites.
S. Levine and V. Koltun, “Guided policy search,” in International Conference on Machine Learning (ICML) , 2013
2013
Later among the works it cites.
——, “Variational policy search via trajectory optimization,” in Advances in Neural Information Processing Systems (NIPS) , 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Theodorou, J. Buchli, and S. Schaal, “Reinforcement learning of motor skills in high dimensions,” in International Conference on Robotics and Automation (ICRA) , 2010
2010
Cited alongside, same era.
J. Peters, K. Mülling, and Y. Altün, “Relative entropy policy search,” in AAAI Conference on Artificial Intelligence , 2010
2010
Cited alongside, same era.
S. M. Khansari-Zadeh and A. Billard, “BM: An iterative algorithm to learn stable non-linear dynamical systems with Gaussian mixture models,” in International Conference on Robotics and Automation (ICRA) , 2010
2010
Cited alongside, same era.
M. Deisenroth, C. Rasmussen, and D. Fox, “Learning to control a low-cost manipulator using data-efficient reinforcement learning,” in Robotics: Science and Systems , 2011
2011
Cited alongside, same era.
M. Kalakrishnan, L. Righetti, P. Pastor, and S. Schaal, “Learning force control policies for compliant manipulation,” in International Conference on Intelligent Robots and Systems (IROS) , 2011
2011
Cited alongside, same era.
P. Pastor, M. Kalakrishnan, S. Chitta, E. Theodorou, and S. Schaal, “Skill learning and task outcome prediction for manipulation,” in International Conference on Robotics and Automation (ICRA) , 2011
2011
Cited alongside, same era.
2013
Later among the works it cites.
S. Levine, “Exploring deep and recurrent architectures for optimal control,” in NIPS 2013 Workshop on Deep Learning , 2013
2013
Later among the works it cites.
S. Levine and P. Abbeel, “Learning neural network policies with guided policy search under unknown dynamics,” in Advances in Neural Information Processing Systems (NIPS) , 2014
2014
Later among the works it cites.
——, “Learning complex neural network policies with trajectory optimization,” in International Conference on Machine Learning (ICML) , 2014
2014
Later among the works it cites.
I. Mordatch and E. Todorov, “Combining the benefits of function approximation and trajectory optimization,” in Robotics: Science and Systems (RSS) , 2014
2014
Later among the works it cites.
R. Lioutikov, A. Paraschos, G. Neumann, and J. Peters, “Sample-based information-theoretic stochastic optimal control,” in International Conference on Robotics and Automation (ICRA) , 2014
2014
Later among the works it cites.