Fetching the paper…
Reading the bibliography…
One of the key challenges in applying reinforcement learning to complex robotic control tasks is the need to gather large amounts of experience in order to find an effective policy for the task at hand.
K. J. Astrom and B. Wittenmark, Adaptive Control , 2nd ed. Boston, MA, USA: Addison-Wesley Longman Publishing Co., Inc., 1994
1994
Earlier work this paper cites.
W. Li and E. Todorov, “Iterative linear quadratic regulator design for nonlinear biological movement systems,” in ICINCO (1) , 2004, pp. 222–229
2004
Earlier work this paper cites.
H. Fukushima, T. Kim, and T. Sugie, “Adaptive model predictive control for a class of constrained linear systems based on the comparison model,” Automatica , vol. 43, no. 2, pp. 301–308, Feb. 2007
2007
Earlier work this paper cites.
J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,” Neural Networks , vol. 21, no. 4, pp. 682–697, 2008
2008
Earlier work this paper cites.
D. Nguyen-Tuong and J. Peters, “Using model knowledge for learning inverse dynamics,” in International Conference on Robotics and Automation (ICRA) , 2010
2010
Earlier work this paper cites.
Y. Wu and Y. Demiris, “Towards one shot learning by imitation for humanoid robots,” in International Conference on Robotics and Automation (ICRA) , 2010
2010
Earlier work this paper cites.
M. Deisenroth and C. Rasmussen, “PILCO: a model-based and data-efficient approach to policy search,” in International Conference on Machine Learning (ICML) , 2011
2011
Earlier work this paper cites.
D. Windate, N. D. Goodman, D. M. Roy, L. P. Kaelbling, and J. B. Tenenbaum, “Bayesian policy search with policy priors,” Proceedings of the Twenty-Second International Joint Conference on Artificial Intelligence (IJCAI) , 2011
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “MuJoCo: A physics engine for model-based control,” in IEEE/RSJ International Conference on Intelligent Robots and Systems , 2012
2012
Earlier work this paper cites.
A. Aswani, P. Bouffard, and C. Tomlin, “Extensions of learning-based model predictive control for real-time application to a quadrotor helicopter,” in American Control Conference (ACC) , 2012
2012
Cited alongside, same era.
Y. Tassa, T. Erez, and E. Todorov, “Synthesis and stabilization of complex behaviors through online trajectory optimization,” in IEEE/RSJ International Conference on Intelligent Robots and Systems , 2012
2012
Cited alongside, same era.
M. Riedmiller, S. Lange, and A. Voigtlaender, “Autonomous reinforcement learning on raw visual input data in a real world application,” in International Joint Conference on Neural Networks , 2012
2012
Cited alongside, same era.
M. Deisenroth, G. Neumann, and J. Peters, “A survey on policy search for robotics,” Foundations and Trends in Robotics , vol. 2, no. 1-2, pp. 1–142, 2013
2013
Cited alongside, same era.
Y. Pan and E. Theodorou, “Probabilistic differential dynamic programming,” in Neural Information Processing Systems (NIPS) , 2014
2014
Later among the works it cites.
M. Cutler and J. P. How, “Efficient reinforcement learning for robots using informative simulated priors,” in International Conference on Robotics and Automation (ICRA) , 2015
2015
Closest in time.
I. Lenz, R. Knepper, and A. Saxena, “DeepMPC: Learning deep latent features for model predictive control,” in Robotics: Science and Systems (RSS) , 2015
2015
Closest in time.
A. Punjani and P. Abbeel, “Deep learning helicopter dynamics models,” in International Conference on Robotics and Automation (ICRA) , 2015
2015
Closest in time.
S. Levine, N. Wagener, and P. Abbeel, “Learning contact-rich manipulation skills with guided policy search,” in International Conference on Robotics and Automation (ICRA) , 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” International Journal of Robotic Research , vol. 32, no. 11, pp. 1238–1274, 2013
2013
Cited alongside, same era.
G. Chowdhary, M. Mühlegg, J. How, and F. Holzapfel, “Concurrent learning adaptive model predictive control,” in Advances in Aerospace Guidance, Navigation and Control . Springer Berlin Heidelberg, 2013
2013
Cited alongside, same era.
S. Levine and V. Koltun, “Guided policy search,” in International Conference on Machine Learning (ICML) , 2013
2013
Cited alongside, same era.
J. Boedecker, J. T. Springenberg, J. Wulfing, and M. Riedmiller, “Approximate real-time optimal control based on sparse Gaussian process models,” in Adaptive Dynamic Programming and Reinforcement Learning (ADPRL) , 2014
2014
Cited alongside, same era.
S. Levine and P. Abbeel, “Learning neural network policies with guided policy search under unknown dynamics,” in Advances in Neural Information Processing Systems (NIPS) , 2014
2014
Cited alongside, same era.
2015
Closest in time.
M. Watter, J. T. Springenberg, J. Boedecker, and M. Riedmiller, “Embed to control: A locally linear latent dynamics model for control from raw images,” in Advances in Neural Information Processing Systems (NIPS) , 2015
2015
Closest in time.
C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel, “Deep spatial autoencoders for visuomotor learning,” International Conference on Robotics and Automation (ICRA) , 2016
2016
Closest in time.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” Journal of Machine Learning Research (JMLR) , 2016
2016
Closest in time.