Fetching the paper…
Reading the bibliography…
We present a model-based framework for robot locomotion that achieves walking based on only 4.5 minutes (45,000 control steps) of data collected on a quadruped robot.
Policy gradient reinforcement learning for fast quadrupedal locomotion
N. Kohl and P. Stone · 2004
Earlier work this paper cites.
The cross-entropy method: A unified approach to monte carlo simulation, randomized optimization and machine learning
R. Y. Rubinstein and D. P. Kroese · 2004
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Model regularization for stable sample rollouts
E. Talvitie · 2014
Earlier work this paper cites.
Optimization-based locomotion planning, estimation, and control design for the atlas humanoid robot
S. Kuindersma, R. Deits, M. Fallon, A. Valenzuela, H. Dai, F. Permenter, T. Koolen, P. Marion, and R. Tedrake · 2016
Earlier work this paper cites.
Design principles for a family of direct-drive legged robots
G. Kenneally, A. De, and D. E. Koditschek · 2016
Earlier work this paper cites.
Anymal-toward legged robots for harsh environments
M. Hutter, C. Gehring, A. Lauber, F. Gunther, C. D. Bellicoso, V. Tsounis, P. Fankhauser, R. Diethelm, S. Bachmann, M. Blösch, et al · 2017
Earlier work this paper cites.
Design of dynamic legged robots
S. Kim, P. M. Wensing, et al · 2017
Earlier work this paper cites.
Emergence of locomotion behaviours in rich environments
N. Heess, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, S. Eslami, M. Riedmiller, et al · 2017
Earlier work this paper cites.
Deeploco: Dynamic locomotion skills using hierarchical deep reinforcement learning
X. B. Peng, G. Berseth, K. Yin, and M. Van De Panne · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Self-correcting models for model-based reinforcement learning
E. Talvitie · 2017
Cited alongside, same era.
Learning locomotion skills using deeprl: Does the choice of action space matter?
X. B. Peng and M. van de Panne · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
MIT cheetah 3: Design and control of a robust, dynamic quadruped robot
G. Bledt, M. J. Powell, B. Katz, J. Di Carlo, P. M. Wensing, and S. Kim · 2018
Cited alongside, same era.
Sim-to-real: Learning agile locomotion for quadruped robots
J. Tan, T. Zhang, E. Coumans, A. Iscen, Y. Bai, D. Hafner, S. Bohez, and V. Vanhoucke · 2018
Cited alongside, same era.
Feedback control for cassie with deep reinforcement learning
Whole-body nonlinear model predictive control through contacts for quadrupeds
M. Neunert, M. Stäuble, M. Giftthaler, C. D. Bellicoso, J. Carius, C. Gehring, M. Hutter, and J. Buchli · 2018
Later among the works it cites.
Dynamic locomotion in the mit cheetah 3 through convex model-predictive control
J. Di Carlo, P. M. Wensing, B. Katz, G. Bledt, and S. Kim · 2018
Later among the works it cites.
Fast online trajectory optimization for the bipedal robot cassie
T. Apgar, P. Clary, K. Green, A. Fern, and J. W. Hurst · 2018
Later among the works it cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, and S. Levine · 2018
Later among the works it cites.
JAX: composable transformations of Python+NumPy programs, 2018
R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, and S. Wanderman-Milne · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Xie, G. Berseth, P. Clary, J. Hurst, and M. van de Panne · 2018
Cited alongside, same era.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2018
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2018
Cited alongside, same era.
Policies modulating trajectory generators
A. Iscen, K. Caluwaerts, J. Tan, T. Zhang, E. Coumans, V. Sindhwani, and V. Vanhoucke · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Cited alongside, same era.
Model-based reinforcement learning via meta-policy optimization
I. Clavera, J. Rothfuss, J. Schulman, Y. Fujita, T. Asfour, and P. Abbeel · 2018
Cited alongside, same era.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Learning agile and dynamic motor skills for legged robots
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter · 2019
Closest in time.
Learning to walk via deep reinforcement learning
T. Haarnoja, S. Ha, A. Zhou, J. Tan, G. Tucker, and S. Levine · 2019
Closest in time.
Autonomous functional movements in a tendon-driven limb via limited experience
A. Marjaninejad, D. Urbina-Meléndez, B. A. Cohn, and F. J. Valero-Cuevas · 2019
Closest in time.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Closest in time.
TF-Agents: A library for reinforcement learning in tensorflow
Sergio Guadarrama, Anoop Korattikara, Oscar Ramirez, Pablo Castro, Ethan Holly, Sam Fishman, Ke Wang, Ekaterina Gonina, Neal Wu, Chris Harris, Vincent Vanhoucke, Eugene Brevdo · 2019
Closest in time.