Fetching the paper…
Reading the bibliography…
We apply multi-agent deep reinforcement learning (RL) to train end-to-end robot soccer policies with fully onboard computation and sensing via egocentric RGB vision.
Iterative reinforcement learning based design of dynamic locomotion skills for Cassie
Z. Xie, P. Clary, J. Dao, P. Morais, J. W. Hurst, and M. van de Panne · 1903
Earlier work this paper cites.
Iterative solution of games by fictitious play
G. W. Brown · 1951
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
A. L. Samuel · 1959
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Finding structure in reinforcement learning
S. Thrun and A. Schwartz · 1994
Earlier work this paper cites.
Temporal difference learning and TD-gammon
G. Tesauro · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
RoboCup: The robot world cup initiative
H. Kitano, M. Asada, Y. Kuniyoshi, I. Noda, and E. Osawa · 1997
Earlier work this paper cites.
Reusing learned policies between similar problems
M. Bowling and M. Veloso · 1998
Earlier work this paper cites.
Karlsruhe Brainstormers - a reinforcement learning approach to robotic soccer
M. Riedmiller, A. Merke, D. Meier, A. Hoffmann, A. Sinner, O. Thate, and R. Ehrmann · 2000
Earlier work this paper cites.
Layered learning
P. Stone and M. Veloso · 2000
Earlier work this paper cites.
Deep Blue
M. Campbell, A. J. H. Jr., and F. Hsu · 2002
Earlier work this paper cites.
Autonomous learning of stable quadruped locomotion
M. Saggar, T. D’Silva, N. Kohl, and P. Stone · 2007
Earlier work this paper cites.
Reinforcement learning for robot soccer
M. Riedmiller, T. Gabel, R. Hafner, and S. Lange · 2009
Earlier work this paper cites.
Learning powerful kicks on the Aibo ERS-7: The quest for a striker
M. Hausknecht and P. Stone · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Fictitious self-play in extensive-form games
J. Heinrich, M. Lanctot, and D. Silver · 2015
Earlier work this paper cites.
Ensemble-CIO: Full-body dynamic motion planning that transfers to physical humanoids
I. Mordatch, K. Lowrey, and E. Todorov · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel · 2017
Earlier work this paper cites.
Matterport3d: Learning from RGB-D data in indoor environments
A. X. Chang, A. Dai, T. A. Funkhouser, M. Halber, M. Nießner, M. Savva, S. Song, A. Zeng, and Y. Zhang · 2017
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
M. Lanctot, V. Zambaldi, A. Gruslys, A. Lazaridou, K. Tuyls, J. Pérolat, D. Silver, and T. Graepel · 2017
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. Riedmiller · 2017
Earlier work this paper cites.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Earlier work this paper cites.
Asymmetric actor critic for image-based robot learning
L. Pinto, M. Andrychowicz, P. Welinder, W. Zaremba, and P. Abbeel · 2018
Earlier work this paper cites.
Gibson env: Real-world perception for embodied agents
F. Xia, A. R. Zamir, Z. He, A. Sax, J. Malik, and S. Savarese · 2018
Earlier work this paper cites.
Overlapping layered learning
P. MacAlpine and P. Stone · 2018
Cited alongside, same era.
Sim-to-real: Learning agile locomotion for quadruped robots, 2018
J. Tan, T. Zhang, E. Coumans, A. Iscen, Y. Bai, D. Hafner, S. Bohez, and V. Vanhoucke · 2018
Cited alongside, same era.
Emergent complexity via multi-agent competition
T. Bansal, J. Pachocki, S. Sidor, I. Sutskever, and I. Mordatch · 2018
Cited alongside, same era.
Kickstarting deep reinforcement learning
S. Schmitt, J. J. Hudson, A. Zídek, S. Osindero, C. Doersch, W. M. Czarnecki, J. Z. Leibo, H. Küttler, A. Zisserman, K. Simonyan, and S. M. A. Eslami · 2018
Cited alongside, same era.
Learning an embedding space for transferable robot skills
K. Hausman, J. T. Springenberg, Z. Wang, N. Heess, and M. Riedmiller · 2018
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Blind bipedal stair traversal via sim-to-real reinforcement learning
J. Siekmann, K. Green, J. Warila, A. Fern, and J. W. Hurst · 2021
Later among the works it cites.
Nerf2real: Sim2real transfer of vision-guided bipedal motion skills using neural radiance fields
A. Byravan, J. Humplik, L. Hasenclever, A. Brussee, F. Nori, T. Haarnoja, B. Moran, S. Bohez, F. Sadeghi, B. Vujatovic, and N. M. O. Heess · 2022
Later among the works it cites.
Learning robust perceptive locomotion for quadrupedal robots in the wild
T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter · 2022
Later among the works it cites.
Robot learning from randomized simulations: A review
F. Muratore, F. Ramos, G. Turk, W. Yu, M. Gienger, and J. Peters · 2022
Later among the works it cites.
S. Masuda and K. Takahashi · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2018
Cited alongside, same era.
Solving rubik’s cube with a robot hand
OpenAI, I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, J. Schneider, N. Tezak, J. Tworek, P. Welinder, L. Weng, Q. Yuan, W. Zaremba, and L. Zhang · 2019
Cited alongside, same era.
Scrutinizing and de-biasing intuitive physics with neural stethoscopes
F. Fuchs, O. Groth, A. Kosiorek, A. Bewley, M. Wulfmeier, A. Vedaldi, and I. Posner · 2019
Cited alongside, same era.
Sim-to-real transfer for biped locomotion
W. Yu, V. C. Kumar, G. Turk, and C. K. Liu · 2019
Cited alongside, same era.
Learning agile and dynamic motor skills for legged robots
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V. Dalibard, D. Budden, Y. Sulsky, J. Molloy, T. L. Paine, Ç. Gülçehre, Z. Wang, T. Pfaff, Y. Wu, R. Ring, D. Yogatama, D. Wünsch, K. McKinney, O. Smith, T. Schaul, T. P. Lillicrap, K. Kavukcuoglu, D. Hassabis, C. Apps, and D. Silver · 2019
Cited alongside, same era.
Emergent coordination through competition
S. Liu, G. Lever, J. Merel, S. Tunyasuvunakool, N. Heess, and T. Graepel · 2019
Cited alongside, same era.
Learning transferable motor skills with hierarchical latent mixture policies
D. Rao, F. Sadeghi, L. Hasenclever, M. Wulfmeier, M. Zambelli, G. Vezzani, D. Tirumala, Y. Aytar, J. Merel, N. Heess, and R. Hadsell · 2022
Later among the works it cites.
Behavior priors for efficient reinforcement learning
D. Tirumala, A. Galashov, H. Noh, L. Hasenclever, R. Pascanu, J. Schwarz, G. Desjardins, W. M. Czarnecki, A. Ahuja, Y. W. Teh, et al · 2022
Later among the works it cites.
Skills: Adaptive skill sequencing for efficient temporally-extended exploration
G. Vezzani, D. Tirumala, M. Wulfmeier, D. Rao, A. Abdolmaleki, B. Moran, T. Haarnoja, J. Humplik, R. Hafner, M. Neunert, et al · 2022
Later among the works it cites.
Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble
S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin · 2022
Later among the works it cites.
Don’t start from scratch: Leveraging prior data to automate robotic reinforcement learning
H. Walke, J. Yang, A. Yu, A. Kumar, J. Orbik, A. Singh, and S. Levine · 2022
Later among the works it cites.
The challenges of exploration for offline reinforcement learning
N. Lambert, M. Wulfmeier, W. F. Whitney, A. Byravan, M. Bloesch, V. Dasagi, T. Hertweck, and M. A. Riedmiller · 2022
Later among the works it cites.
Cassie sets world record for 100m run, 2022
Agility Robotics · 2022
Later among the works it cites.
Towards real robot learning in the wild: A case study in bipedal locomotion
M. Bloesch, J. Humplik, V. Patraucean, R. Hafner, T. Haarnoja, A. Byravan, N. Y. Siegel, S. Tunyasuvunakool, F. Casarini, N. Batchelor, et al · 2022
Later among the works it cites.
Extreme parkour with legged robots
X. Cheng, K. Shi, A. Agarwal, and D. Pathak · 2023
Later among the works it cites.
Robot parkour learning
Z. Zhuang, Z. Fu, J. Wang, C. G. Atkeson, S. Schwertfeger, C. Finn, and H. Zhao · 2023
Later among the works it cites.
Champion-level drone racing using deep reinforcement learning
E. Kaufmann, L. Bauersfeld, A. Loquercio, M. Müller, V. Koltun, and D. Scaramuzza · 2023
Later among the works it cites.
Robotis OP3 manual
Robotis · 2023
Later among the works it cites.
Dribblebot: Dynamic legged manipulation in the wild
Y. Ji, G. B. Margolis, and P. Agrawal · 2023
Later among the works it cites.
Learning to look by self-prediction
M. K. Grimes, J. V. Modayil, P. W. Mirowski, D. Rao, and R. Hadsell · 2023
Later among the works it cites.
Foundations for transfer in reinforcement learning: A taxonomy of knowledge modalities
M. Wulfmeier, A. Byravan, S. Bechtle, K. Hausman, and N. M. O. Heess · 2023
Later among the works it cites.
Robust and versatile bipedal jumping control through multi-task reinforcement learning
Z. Li, X. B. Peng, P. Abbeel, S. Levine, G. Berseth, and K. Sreenath · 2023
Later among the works it cites.
Anymal parkour: Learning agile navigation for quadrupedal robots
D. Hoeller, N. Rudin, D. V. Sako, and M. Hutter · 2024
Closest in time.
Learning agile soccer skills for a bipedal robot with deep reinforcement learning
T. Haarnoja, B. Moran, G. Lever, S. H. Huang, D. Tirumala, J. Humplik, M. Wulfmeier, S. Tunyasuvunakool, N. Y. Siegel, R. Hafner, M. Bloesch, K. Hartikainen, A. Byravan, L. Hasenclever, Y. Tassa, F. Sadeghi, N. Batchelor, F. Casarini, S. Saliceti, C. Game, N. Sreendra, K. Patel, M. Gwira, A. Huber, N. Hurley, F. Nori, R. Hadsell, and N. Heess · 2024
Closest in time.
Real-world humanoid locomotion with reinforcement learning
I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath · 2024
Closest in time.
Replay across experiments: A natural extension of off-policy RL
D. Tirumala, T. Lampe, J. E. Chen, T. Haarnoja, S. Huang, G. Lever, B. Moran, T. Hertweck, L. Hasenclever, M. Riedmiller, N. Heess, and M. Wulfmeier · 2024
Closest in time.
Robocup project
RoboCup Federation · 2024
Closest in time.