Fetching the paper…
Reading the bibliography…
Reinforcement learning has emerged as a promising methodology for training robot controllers.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, no. 3, pp. 229–256, 1992
1992
Earlier work this paper cites.
L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” Journal of artificial intelligence research , vol. 4, pp. 237–285, 1996
1996
Earlier work this paper cites.
S. Amari, “Natural gradient works efficiently in learning,” Neural Computation , vol. 10, pp. 251–276, 1998
1998
Earlier work this paper cites.
S. Kakade, “A natural policy gradient,” in NIPS , 2001
2001
Earlier work this paper cites.
J. Peters, S. Vijayakumar, and S. Schaal, “Reinforcement learning for humanoid robotics,” in Proceedings of the third IEEE-RAS international conference on humanoid robots , 2003, pp. 1–20
2003
Earlier work this paper cites.
I. Menache, S. Mannor, and N. Shimkin, “Basis function adaptation in temporal difference reinforcement learning,” Annals of Operations Research , vol. 134, no. 1, pp. 215–238, 2005
2005
Earlier work this paper cites.
P. Abbeel, M. Quigley, and A. Y. Ng, “Using inaccurate models in reinforcement learning,” in Proceedings of the 23rd international conference on Machine learning . ACM, 2006, pp. 1–8
2006
Earlier work this paper cites.
J. Peters and S. Schaal, “Natural actor-critic,” Neurocomputing , vol. 71, pp. 1180–1190, 2007
2007
Earlier work this paper cites.
J. Peters, “Machine learning of motor skills for robotics,” PhD Dissertation, University of Southern California , 2007
2007
Earlier work this paper cites.
M. Raibert, K. Blankespoor, G. Nelson, and R. Playter, “Bigdog, the rough-terrain quadruped robot,” IFAC Proceedings Volumes , vol. 41, no. 2, pp. 10 822–10 825, 2008
2008
Earlier work this paper cites.
J. Kober and J. R. Peters, “Policy search for motor primitives in robotics,” in Advances in neural information processing systems , 2009, pp. 849–856
2009
Earlier work this paper cites.
J. M. Wang, D. J. Fleet, and A. Hertzmann, “Optimizing walking controllers for uncertain inputs and environments,” in ACM Transactions on Graphics (TOG) , vol. 29, no. 4. ACM, 2010, p. 73
2010
Earlier work this paper cites.
S. Barrett, M. E. Taylor, and P. Stone, “Transfer learning for reinforcement learning on a physical robot,” in Ninth International Conference on Autonomous Agents and Multiagent Systems-Adaptive Learning Agents Workshop (AAMAS-ALA) , 2010
2010
Earlier work this paper cites.
D. Nguyen-Tuong and J. Peters, “Model learning for robot control: a survey,” Cognitive processing , vol. 12, no. 4, pp. 319–340, 2011
2011
Earlier work this paper cites.
M. Deisenroth and C. E. Rasmussen, “Pilco: A model-based and data-efficient approach to policy search,” in Proceedings of the 28th International Conference on machine learning (ICML-11) , 2011, pp. 465–472
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” pp. 5026–5033, 2012
2012
Cited alongside, same era.
Y. Tassa, T. Erez, and E. Todorov, “Synthesis and stabilization of complex behaviors through online trajectory optimization,” in Intelligent Robots and Systems (IROS), 2012 IEEE/RSJ International Conference on . IEEE, 2012, pp. 4906–4913
2012
Cited alongside, same era.
I. Mordatch, E. Todorov, and Z. Popović, “Discovery of complex behaviors through contact-invariant optimization,” ACM Transactions on Graphics (TOG) , vol. 31, no. 4, p. 43, 2012
2012
Cited alongside, same era.
2012
Cited alongside, same era.
J. Schulman, S. Levine, P. Moritz, M. Jordan, and P. Abbeel, “Trust region policy optimization,” in ICML , 2015
2015
Later among the works it cites.
I. Mordatch, K. Lowrey, G. Andrew, Z. Popovic, and E. V. Todorov, “Interactive control of diverse complex characters with neural networks,” in Advances in Neural Information Processing Systems , 2015, pp. 3132–3140
2015
Later among the works it cites.
2016
Later among the works it cites.
I. Mordatch, N. Mishra, C. Eppner, and P. Abbeel, “Combining model-based policy search with online model learning for control of physical humanoids,” in Robotics and Automation (ICRA), 2016 IEEE International Conference on . IEEE, 2016, pp. 242–248
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2012
Cited alongside, same era.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research , vol. 32, no. 11, pp. 1238–1274, 2013
2013
Cited alongside, same era.
Y. Tassa, N. Mansard, and E. Todorov, “Control-limited differential dynamic programming,” in Robotics and Automation (ICRA), 2014 IEEE International Conference on . IEEE, 2014, pp. 1168–1175
2014
Cited alongside, same era.
M. Cutler, T. J. Walsh, and J. P. How, “Reinforcement learning with multi-fidelity simulators,” in Robotics and Automation (ICRA), 2014 IEEE International Conference on . IEEE, 2014, pp. 3888–3895
2014
Cited alongside, same era.
S. Kolev and E. Todorov, “Physically consistent state estimation and system identification for contacts,” in Humanoid Robots (Humanoids), 2015 IEEE-RAS 15th International Conference on . IEEE, 2015, pp. 1036–1043
2015
Cited alongside, same era.
I. Mordatch, K. Lowrey, and E. Todorov, “Ensemble-cio: Full-body dynamic motion planning that transfers to physical humanoids,” in Intelligent Robots and Systems (IROS), 2015 IEEE/RSJ International Conference on . IEEE, 2015, pp. 5307–5314
2015
Cited alongside, same era.
S. Feng, E. Whitman, X. Xinjilefu, and C. G. Atkeson, “Optimization-based full body control for the darpa robotics challenge,” Journal of Field Robotics , vol. 32, no. 2, pp. 293–312, 2015
2015
Cited alongside, same era.
J. Koenemann, A. Del Prete, Y. Tassa, E. Todorov, O. Stasse, M. Bennewitz, and N. Mansard, “Whole-body model-predictive control applied to the hrp-2 humanoid,” in Intelligent Robots and Systems (IROS), 2015 IEEE/RSJ International Conference on . IEEE, 2015, pp. 3346–3351
2015
Cited alongside, same era.
2016
Later among the works it cites.
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, “High-dimensional continuous control using generalized advantage estimation,” in ICLR , 2016
2016
Later among the works it cites.
A. Rajeswaran, K. Lowrey, E. Todorov, and S. Kakade, “Towards Generalization and Simplicity in Continuous Control,” in NIPS , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Gu, E. Holly, T. Lillicrap, and S. Levine, “Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates,” in Robotics and Automation (ICRA), 2017 IEEE International Conference on . IEEE, 2017, pp. 3389–3396
2017
Later among the works it cites.
Y. Chebotar, M. Kalakrishnan, A. Yahya, A. Li, S. Schaal, and S. Levine, “Path integral guided policy search,” in Robotics and Automation (ICRA), 2017 IEEE International Conference on . IEEE, 2017, pp. 3381–3388
2017
Later among the works it cites.
2017
Later among the works it cites.
W. Montgomery, A. Ajay, C. Finn, P. Abbeel, and S. Levine, “Reset-free guided policy search: efficient deep reinforcement learning with stochastic initial states,” in Robotics and Automation (ICRA), 2017 IEEE International Conference on . IEEE, 2017, pp. 3373–3380
2017
Later among the works it cites.