R. Tedrake, T. W. Zhang, and H. S. Seung, “Learning to walk in 20 minutes,” in Proceedings of the Fourteenth Yale Workshop on Adaptive and Learning Systems , vol. 95585. Yale University New Haven (CT), 2005, pp. 1939–1412
1939
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Proceedings of the 12th International Conference on Neural Information Processing Systems , ser. NIPS’99. Cambridge, MA, USA: MIT Press, 1999, pp. 1057–1063. [Online]. Available: http://dl.acm.org/citation.cfm?id=3009657.3009806
1999
Earlier work this paper cites.
E. Schuitema, M. Wisse, T. Ramakers, and P. Jonker, “The design of leo: A 2d bipedal walking robot for online autonomous reinforcement learning,” in 2010 IEEE/RSJ International Conference on Intelligent Robots and Systems , Oct 2010, pp. 3238–3243
2010
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , Oct 2012, pp. 5026–5033
2012
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” CoRR , vol. abs/1412.6980, 2014. [Online]. Available: http://arxiv.org/abs/1412.6980
Original
2014
Earlier work this paper cites.
S. C. Hsu, X. Xu, and A. D. Ames, “Control barrier function based quadratic programs with application to bipedal robotic walking,” in 2015 American Control Conference (ACC) , July 2015, pp. 4542–4548
2015
Earlier work this paper cites.
M. Posa, S. Kuindersma, and R. Tedrake, “Optimization and stabilization of trajectories for constrained dynamical systems,” in 2016 IEEE International Conference on Robotics and Automation (ICRA) , May 2016, pp. 1366–1373
2016
Earlier work this paper cites.