Fetching the paper…
Reading the bibliography…
We investigate whether Deep Reinforcement Learning (Deep RL) is able to synthesize sophisticated and safe movement skills for a low-cost, miniature humanoid robot that can be composed into complex behavioral strategies in dynamic environments.
G. W. Brown, “Iterative solution of games by fictitious play,” in Activity Analysis of Production and Allocation
1951
Earlier work this paper cites.
MIT press, 1986
M. H. Raibert, Legged robots that balance · 1986
Earlier work this paper cites.
K. Sims, “Evolving virtual creatures,” in Proceedings of the 21st Annual Conference on Computer Graphics and Interactive Techniques
1994
Earlier work this paper cites.
S. Thrun and A. Schwartz, “Finding structure in reinforcement learning,” Advances in neural information processing systems
1994
Earlier work this paper cites.
H. Kitano, M. Asada, Y. Kuniyoshi, I. Noda, and E. Osawa, “RoboCup: The robot world cup initiative,” in Proceedings of the first international conference on Autonomous agents
1997
Earlier work this paper cites.
M. Bowling and M. Veloso, “Reusing learned policies between similar problems,” in Proceedings of the AI* AI-98 Workshop on New Trends in Robotics
1998
Earlier work this paper cites.
R. S. Sutton, D. Precup, and S. Singh, “Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,” Artificial intelligence
1999
Earlier work this paper cites.
M. Riedmiller, A. Merke, D. Meier, A. Hoffmann, A. Sinner, O. Thate, and R. Ehrmann, “Karlsruhe Brainstormers - a reinforcement learning approach to robotic soccer,” in RoboCup-2000: Robot Soccer World Cup IV, LNCS
2000
Earlier work this paper cites.
P. Stone and M. Veloso, “Layered learning,” in European conference on machine learning
2000
Earlier work this paper cites.
K. Tuyls, S. Maes, and B. Manderick, “Reinforcement learning in large state spaces,” in RoboCup 2002: Robot Soccer World Cup VI
2002
Earlier work this paper cites.
P. Stone, R. S. Sutton, and G. Kuhlmann, “Reinforcement learning for RoboCup-soccer keepaway,” Adaptive Behavior
2005
Earlier work this paper cites.
S. Kalyanakrishnan, Y. Liu, and P. Stone, “Half field offense in RoboCup soccer: A multiagent reinforcement learning case study,” in RoboCup-2006: Robot Soccer World Cup X
2007
Earlier work this paper cites.
M. Saggar, T. D’Silva, N. Kohl, and P. Stone, “Autonomous learning of stable quadruped locomotion,” in RoboCup-2006: Robot Soccer World Cup X
2007
Earlier work this paper cites.
J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,” Neural networks
2008
Earlier work this paper cites.
M. Riedmiller, T. Gabel, R. Hafner, and S. Lange, “Reinforcement learning for robot soccer,” Autonomous Robots
2009
Earlier work this paper cites.
S. Kalyanakrishnan and P. Stone, “Learning complementary multiagent behaviors: A case study,” in RoboCup 2009: Robot Soccer World Cup XIII
2010
Earlier work this paper cites.
M. Hausknecht and P. Stone, “Learning powerful kicks on the Aibo ERS-7: The quest for a striker,” in RoboCup-2010: Robot Soccer World Cup XIV
2011
Earlier work this paper cites.
S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in Proceedings of the fourteenth international conference on artificial intelligence and statistics
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems
2012
Earlier work this paper cites.
M. P. Deisenroth, G. Neumann, J. Peters, et al
2013
Earlier work this paper cites.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research
2013
Earlier work this paper cites.
A. Farchy, S. Barrett, P. MacAlpine, and P. Stone, “Humanoid robots learning to walk faster: From the real world to simulation and back,” in Proc. of 12th Int. Conf. on Autonomous Agents and Multiagent Systems (AAMAS)
2013
Earlier work this paper cites.
D. M. Roijers, P. Vamplew, S. Whiteson, and R. Dazeley, “A survey of multi-objective sequential decision-making,” Journal of Artificial Intelligence Research
2013
Earlier work this paper cites.
I. Mordatch, K. Lowrey, and E. Todorov, “Ensemble-CIO: Full-body dynamic motion planning that transfers to physical humanoids,” in 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
2015
Earlier work this paper cites.
J. Heinrich, M. Lanctot, and D. Silver, “Fictitious self-play in extensive-form games,” in Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in Proceedings of the 32nd International Conference on Machine Learning (ICML)
2015
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Earlier work this paper cites.
S. Kuindersma, R. Deits, M. Fallon, A. Valenzuela, H. Dai, F. Permenter, T. Koolen, P. Marion, and R. Tedrake, “Optimization-based locomotion planning, estimation, and control design for the atlas humanoid robot,” Autonomous Robots
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. G. Bellemare, W. Dabney, and R. Munos, “A distributional perspective on reinforcement learning,” in Proceedings of the 34th International Conference on Machine Learning
2017
Earlier work this paper cites.
M. Lanctot, V. Zambaldi, A. Gruslys, A. Lazaridou, K. Tuyls, J. Pérolat, D. Silver, and T. Graepel, “A unified game-theoretic approach to multiagent reinforcement learning,” in Advances in neural information processing systems
2017
Earlier work this paper cites.
Y. Teh, V. Bapst, W. M. Czarnecki, J. Quan, J. Kirkpatrick, R. Hadsell, N. Heess, and R. Pascanu, “Distral: Robust multitask reinforcement learning,” in Advances in Neural Information Processing Systems
2017
Earlier work this paper cites.
T. Bansal, J. Pachocki, S. Sidor, I. Sutskever, and I. Mordatch, “Emergent complexity via multi-agent competition,” in 6th International Conference on Learning Representations (ICLR)
2018
Earlier work this paper cites.
X. B. Peng, P. Abbeel, S. Levine, and M. van de Panne, “DeepMimic: Example-guided deep reinforcement learning of physics-based character skills,” ACM Transactions on Graphics (TOG)
2018
Earlier work this paper cites.
L. McInnes, J. Healy, and J. Melville, “UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction,” ArXiv e-prints
2018
Earlier work this paper cites.
P. MacAlpine and P. Stone, “Overlapping layered learning,” Artificial Intelligence
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller, “Maximum a posteriori policy optimisation,” in Proceedings of the 6th International Conference on Learning Representations (ICLR)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
M. Riedmiller, R. Hafner, T. Lampe, M. Neunert, J. Degrave, T. van de Wiele, V. Mnih, N. Heess, and J. T. Springenberg, “Learning by playing solving sparse reward tasks from scratch,” in Proceedings of the 35th International Conference on Machine Learning
2018
R. Hafner, T. Hertweck, P. Klöppner, M. Bloesch, M. Neunert, M. Wulfmeier, S. Tunyasuvunakool, N. Heess, and M. Riedmiller, “Towards general and autonomous learning of core skills: A case study in locomotion,” in Conference on Robot Learning
2021
Later among the works it cites.
S. Liu, G. Lever, Z. Wang, J. Merel, S. M. A. Eslami, D. Hennes, W. M. Czarnecki, Y. Tassa, S. Omidshafiei, A. Abdolmaleki, N. Y. Siegel, L. Hasenclever, L. Marris, S. Tunyasuvunakool, H. F. Song, M. Wulfmeier, P. Muller, T. Haarnoja, B. D. Tracey, K. Tuyls, T. Graepel, and N. Heess, “From motor control to team play in simulated humanoid football,” Science Robotics
2022
Later among the works it cites.
N. Rudin, D. Hoeller, M. Bjelonic, and M. Hutter, “Advanced skills by learning locomotion and local navigation end-to-end,” in 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cambridge, MA, USA: A Bradford Book, 2018
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction · 2018
Cited alongside, same era.
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter, “Learning agile and dynamic motor skills for legged robots,” Science Robotics
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
B. D. DeAngelis, J. A. Zavatone-Veth, and D. A. Clark, “The manifold structure of limb coordination in walking Drosophila
2019
Cited alongside, same era.
Only available online: http://www.b-human.de/downloads/publications/2019/CodeRelease2019.pdf
T. Röfer, T. Laue, A. Baude, J. Blumenkamp, G. Felsch, J. Fiedler, A. Hasselbring, T. Haß, J. Oppermann, P. Reichenberg, N. Schrader, and D. Weiß, “B-Human team report and code release 2019,” 2019 · 2019
Cited alongside, same era.
T. Haarnoja, S. Ha, A. Zhou, J. Tan, G. Tucker, and S. Levine, “Learning to walk via deep reinforcement learning,” in Proceedings of Robotics: Science and Systems (RSS)
2019
Cited alongside, same era.
W. Yu, V. C. Kumar, G. Turk, and C. K. Liu, “Sim-to-real transfer for biped locomotion,” in 2019 ieee/rsj international conference on intelligent robots and systems (IROS)
2019
Cited alongside, same era.
2022
Later among the works it cites.
Y. Ji, Z. Li, Y. Sun, X. B. Peng, S. Levine, G. Berseth, and K. Sreenath, “Hierarchical reinforcement learning for precise soccer shooting skills using a quadrupedal robot,” in 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
2022
Later among the works it cites.
2022
Later among the works it cites.
https://www.youtube.com/watch?v=DdojWYOK0Nc
Agility Robotics, “Cassie sets world record for 100m run,” 2022 · 2022
Later among the works it cites.
RoboCup Federation, “Robocup project.” https://www.robocup.org , May 2022
2022
Later among the works it cites.
M. Bestmann and J. Zhang, “Bipedal walking on humanoid robots through parameter optimization,” in RoboCup 2022: - Robot World Cup XXV [Bangkok, Thailand, July 11-17, 2022]
2022
Later among the works it cites.
T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter, “Learning robust perceptive locomotion for quadrupedal robots in the wild,” Science Robotics
2022
Later among the works it cites.
F. Muratore, F. Ramos, G. Turk, W. Yu, M. Gienger, and J. Peters, “Robot learning from randomized simulations: A review,” Frontiers in Robotics and AI
2022
Later among the works it cites.
2022
Later among the works it cites.
M. Bloesch, J. Humplik, V. Patraucean, R. Hafner, T. Haarnoja, A. Byravan, N. Y. Siegel, S. Tunyasuvunakool, F. Casarini, N. Batchelor, et al
2022
Later among the works it cites.
G. Ji, J. Mun, H. Kim, and J. Hwangbo, “Concurrent training of a control policy and a state estimator for dynamic and robust legged locomotion,” IEEE Robotics and Automation Letters
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Jin, X. Liu, Y. Shao, H. Wang, and W. Yang, “High-speed quadrupedal locomotion by imitation-relaxation reinforcement learning,” Nature Machine Intelligence
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Salter, M. Wulfmeier, D. Tirumala, N. Heess, M. Riedmiller, R. Hadsell, and D. Rao, “Mo2: Model-based offline options,” in Conference on Lifelong Learning Agents
2022
Later among the works it cites.
D. Tirumala, A. Galashov, H. Noh, L. Hasenclever, R. Pascanu, J. Schwarz, G. Desjardins, W. M. Czarnecki, A. Ahuja, Y. W. Teh, et al
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Choi, G. Ji, J. Park, H. Kim, J. Mun, J. H. Lee, and J. Hwangbo, “Learning quadrupedal locomotion on deformable terrain,” Science Robotics
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
https://www.bostondynamics.com/resources/blog/picking-momentum
R. Deits and T. Koolen, “Picking up momentum,” Web blog post, Boston Dynamics · 2023
Closest in time.
Robotis, “Robotis OP3 manual.” https://emanual.robotis.com/docs/en/platform/op3/introduction , March 2023
2023
Closest in time.
Robotis, “Robotis OP3 source code.” https://github.com/ROBOTIS-GIT/ROBOTIS-OP3 , April 2023
2023
Closest in time.
A. Agarwal, A. Kumar, J. Malik, and D. Pathak, “Legged locomotion in challenging terrains using egocentric vision,” in Conference on Robot Learning
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
P. Wu, A. Escontrela, D. Hafner, P. Abbeel, and K. Goldberg, “DayDreamer: World models for physical robot learning,” in Conference on Robot Learning
2023
Closest in time.
2023
Closest in time.
A. Byravan, J. Humplik, L. Hasenclever, A. Brussee, F. Nori, T. Haarnoja, B. Moran, S. Bohez, F. Sadeghi, B. Vujatovic, and N. Heess, “NeRF2Real: Sim2real transfer of vision-guided bipedal motion skills using neural radiance fields,” in Proceedings of IEEE International Conference on Robotics and Automation (ICRA)
2023
Closest in time.
Optitrack, “Motive optical motion capture software.” https://optitrack.com/software/motive/ , March 2023
2023
Closest in time.
2023
Closest in time.
https://doi.org/10.5281/zenodo.10793725
T. Haarnoja, B. Moran, G. Lever, S. H. Huang, D. Tirumala, J. Humplik, M. Wulfmeier, S. Tunyasuvunakool, N. Y. Siegel, R. Hafner, M. Bloesch, K. Hartikainen, A. Byravan, L. Hasenclever, T. Y., F. Sadeghi, N. Batchelor, F. Casarini, S. Saliceti, C. Game, N. Sreendra, K. Patel, M. Gwira, A. Huber, N. Hurley, F. Nori, R. Hadsell, and N. Heess, “Data release for: Learning agile soccer skills for a bipedal robot with deep reinforcement learning [data set].,” 2024 · 2024
Closest in time.