Fetching the paper…
Reading the bibliography…
Reinforcement learning provides a powerful and flexible framework for automated acquisition of robotic motion skills.
B. Espiau, F. Chaumette, and P. Rives, “A new approach to visual servoing in robotics,” IEEE Transactions on Robotics and Automation , vol. 8, no. 3, 1992
1992
Earlier work this paper cites.
B. H. Yoshimi and P. K. Allen, “Active, uncalibrated visual servoing,” in International Conference on Robotics and Automation (ICRA) , 1994
1994
Earlier work this paper cites.
L. Baird et al. , “Residual algorithms: Reinforcement learning with function approximation,” in Proceedings of the twelfth international conference on machine learning , 1995, pp. 30–37
1995
Earlier work this paper cites.
W. J. Wilson, C. W. W. Hulls, and G. S. Bell, “Relative end-effector control using cartesian position based visual servoing,” IEEE Transactions on Robotics and Automation , vol. 12, no. 5, 1996
1996
Earlier work this paper cites.
M. Jägersand, O. Fuentes, and R. C. Nelson, “Experimental evaluation of uncalibrated visual servoing for precision manipulation,” in International Conference on Robotics and Automation (ICRA) , 1997
1997
Earlier work this paper cites.
S. Singh, M. L. Littman, N. K. Jong, D. Pardoe, and P. Stone, “Learning predictive state representations,” in ICML , 2003
2003
Earlier work this paper cites.
J. A. Bagnell and J. Schneider, “Covariant policy search,” in International Joint Conference on Artificial Intelligence (IJCAI) , 2003
2003
Earlier work this paper cites.
N. Kohl and P. Stone, “Policy gradient reinforcement learning for fast quadrupedal locomotion,” in International Conference on Robotics and Automation (IROS) , 2004
2004
Earlier work this paper cites.
R. Tedrake, T. Zhang, and H. Seung, “Stochastic policy gradient reinforcement learning on a simple 3d biped,” in International Conference on Intelligent Robots and Systems (IROS) , 2004
2004
Earlier work this paper cites.
B. Sallans and G. E. Hinton, “Reinforcement learning with factored states and actions,” The Journal of Machine Learning Research , vol. 5, pp. 1063–1088, 2004
2004
Earlier work this paper cites.
W. Li and E. Todorov, “Iterative linear quadratic regulator design for nonlinear biological movement systems,” in ICINCO (1) , 2004
2004
Earlier work this paper cites.
T. Geng, B. Porr, and F. Wörgötter, “Fast biped walking with a reflexive controller and realtime policy searching,” in Advances in Neural Information Processing Systems (NIPS) , 2006
2006
Earlier work this paper cites.
H. Lee, C. Ekanadham, and A. Y. Ng, “Sparse deep belief net model for visual area V2,” in Neural Information Processing Systems (NIPS) , 2008
2008
Earlier work this paper cites.
G. Endo, J. Morimoto, T. Matsubara, J. Nakanishi, and G. Cheng, “Learning CPG-based biped locomotion with a policy gradient method: Application to a humanoid robot,” International Journal of Robotic Research , vol. 27, no. 2, pp. 213–228, 2008
2008
Cited alongside, same era.
J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,” Neural Networks , vol. 21, no. 4, pp. 682–697, 2008
2008
Cited alongside, same era.
P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal, “Learning and generalization of motor skills by learning from demonstration,” in International Conference on Robotics and Automation (ICRA) , 2009
2009
Cited alongside, same era.
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition (CVPR) , 2009
2009
Cited alongside, same era.
R. Jonschkowski and O. Brock, “State representation learning in robotics: Using prior knowledge about physical interaction,” in Robotics: Science and Systems (RSS) , 2014
2014
Later among the works it cites.
S. Levine and P. Abbeel, “Learning neural network policies with guided policy search under unknown dynamics,” in Advances in Neural Information Processing Systems (NIPS) , 2014
2014
Later among the works it cites.
K. Mohta, V. Kumar, and K. Daniilidis, “Vision based control of a quadrotor for perching on planes and lines,” in International Conference on Robotics and Automation (ICRA) , 2014
2014
Later among the works it cites.
2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Kober, E. Oztop, and J. Peters, “Reinforcement learning to adjust robot movements to new situations,” in Robotics: Science and Systems (RSS) , 2010
2010
Cited alongside, same era.
J. Peters, K. Mülling, and Y. Altün, “Relative entropy policy search,” in AAAI Conference on Artificial Intelligence , 2010
2010
Cited alongside, same era.
M. Deisenroth, C. Rasmussen, and D. Fox, “Learning to control a low-cost manipulator using data-efficient reinforcement learning,” in Robotics: Science and Systems (RSS) , 2011
2011
Cited alongside, same era.
B. Boots, S. M. Siddiqi, and G. J. Gordon, “Closing the learning-planning loop with predictive state representations,” The International Journal of Robotics Research , vol. 30, no. 7, pp. 954–966, 2011
2011
Cited alongside, same era.
S. Lange, A. Voigtlaender, and M. Riedmiller, “Autonomous reinforcement learning on raw visual input data in a real world application,” in International Joint Conference on Neural Networks , 2012
2012
Cited alongside, same era.
S. Ross and A. Bagnell, “Agnostic system identification for model-based reinforcement learning,” in International Conference on Machine Learning (ICML) , 2012
2012
Cited alongside, same era.
T. Lampe and M. Riedmiller, “Acquiring visual servoing reaching and grasping skills using neural reinforcement learning,” in International Joint Conference on Neural Networks (IJCNN) , 2013
2013
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing Atari with deep reinforcement learning,” NIPS ’13 Workshop on Deep Learning , 2013
2013
Cited alongside, same era.
M. Watter, J. T. Springenberg, J. Boedecker, and M. Riedmiller, “Embed to control: A locally linear latent dynamics model for control from raw images,” in Advances in Neural Information Processing Systems (NIPS) , 2015
2015
Closest in time.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, pp. 436–444, 2015
2015
Closest in time.
S. Levine, N. Wagener, and P. Abbeel, “Learning contact-rich manipulation skills with guided policy search,” in International Conference on Robotics and Automation (ICRA) , 2015
2015
Closest in time.
W. Böhmer, J. T. Springenberg, J. Boedecker, M. Riedmiller, and K. Obermayer, “Autonomous learning of state representations for control,” KI - Künstliche Intelligenz , pp. 1–10, 2015
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International Conference on Machine Learning, ICML 2015 , 2015, pp. 448–456
2015
Closest in time.