Fetching the paper…
Reading the bibliography…
Intelligent behaviour in the physical world exhibits structure at multiple spatial and temporal scales.
Joel Z. Leibo, Edward Hughes, Marc Lanctot, and Thore Graepel · 1903
Earlier work this paper cites.
Iterative reinforcement learning based design of dynamic locomotion skills for Cassie
Zhaoming Xie, Patrick Clary, Jeremy Dao, Pedro Morais, Jonathan W. Hurst, and Michiel van de Panne · 1903
Earlier work this paper cites.
Google research football: A novel reinforcement learning environment
Karol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac, Olivier Bachem, Lasse Espeholt, Carlos Riquelme, Damien Vincent, Marcin Michalski, Olivier Bousquet, and Sylvain Gelly · 1907
Earlier work this paper cites.
Solving Rubik’s cube with a robot hand
OpenAI, Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, Jonas Schneider, Nikolas Tezak, Jerry Tworek, Peter Welinder, Lilian Weng, Qiming Yuan, Wojciech Zaremba, and Lei Zhang · 1910
Earlier work this paper cites.
The problem of serial order in behavior , volume 21
Karl Spencer Lashley · 1951
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
A. L. Samuel · 1959
Earlier work this paper cites.
Scripts, plans, goals, and understanding: An inquiry into human knowledge structures
Roger C Schank and Robert P Abelson · 1977
Earlier work this paper cites.
The Rating of Chessplayers, Past and Present
Arpad E. Elo · 1978
Earlier work this paper cites.
Arms races between and within species
Richard Dawkins, John Richard Krebs, J. Maynard Smith, and Robin Holliday · 1979
Earlier work this paper cites.
Legged robots that balance
Marc H Raibert · 1986
Earlier work this paper cites.
A robust layered control system for a mobile robot
Rodney Brooks · 1986
Earlier work this paper cites.
Hierarchical organization of motor programs
David A Rosenbaum · 1987
Earlier work this paper cites.
Algorithmic motion planning in robotics
Micha Sharir · 1989
Earlier work this paper cites.
Unified Theories of Cognition
Allen Newell · 1990
Earlier work this paper cites.
Neural sequence chunkers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Combining specialized reasoners and general purpose planners: A case study
Subbarao Kambhampati, Mark Cutkosky, Marty Tenenbaum, and Soo Hong Lee · 1991
Earlier work this paper cites.
A Reference Model Architecture for Intelligent Systems Design , page 27–56
James S. Albus · 1993
Earlier work this paper cites.
The role of deliberate practice in the acquisition of expert performance
K Anders Ericsson, Ralf T Krampe, and Clemens Tesch-Römer · 1993
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1993
Earlier work this paper cites.
Evolving virtual creatures
Karl Sims · 1994
Earlier work this paper cites.
Temporal difference learning and TD-gammon
G. Tesauro · 1995
Earlier work this paper cites.
RoboCup: The robot World Cup initiative
Hiroaki Kitano, Minoru Asada, Yasuo Kuniyoshi, Itsuki Noda, and Eiichi Osawa · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Hq-learning
Marco Wiering and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Ronald Parr and Stuart J Russell · 1998
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition, 1999
Thomas G. Dietterich · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Team-partitioned, opaque-transition reinforcement learning
Peter Stone and Manuela Veloso · 1999
Earlier work this paper cites.
Layered learning in multiagent systems - a winning approach to robotic soccer
Peter Stone · 2000
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
Martin Lauer and Martin A. Riedmiller · 2000
Earlier work this paper cites.
The complexity of decentralized control of Markov Decision Processes
Daniel S. Bernstein, Shlomo Zilberstein, and Neil Immerman · 2000
Earlier work this paper cites.
Karlsruhe Brainstormers - a reinforcement learning approach to robotic soccer
M. Riedmiller, A. Merke, D. Meier, A. Hoffmann, A. Sinner, O. Thate, and R. Ehrmann · 2000
Earlier work this paper cites.
The prefrontal cortex–an update: time is of the essence
JM Fuster · 2001
Earlier work this paper cites.
Composable controllers for physics-based character animation
Petros Faloutsos, Michiel van de Panne, and Demetri Terzopoulos · 2001
Earlier work this paper cites.
Deep Blue
Murray Campbell, A. Joseph Hoane Jr., and Feng-hsiung Hsu · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Andrew Barto and Sridhar Mahadevan · 2002
Earlier work this paper cites.
Reinforcement learning in large state spaces
Karl Tuyls, Sam Maes, and Bernard Manderick · 2002
Earlier work this paper cites.
Invariant visual representation by single neurons in the human brain
Quian R. Quiroga, L. Reddy, G. Kreiman, C. Koch, and I. Fried · 2005
Earlier work this paper cites.
Cooperative multi-agent learning: The state of the art
Liviu Panait and Sean Luke · 2005
Earlier work this paper cites.
Reinforcement learning for RoboCup-soccer keepaway
Peter Stone, Richard S. Sutton, and Gregory Kuhlmann · 2005
Earlier work this paper cites.
Joint action: bodies and minds moving together
Natalie Sebanz, Harold Bekkering, and Günther Knoblich · 2006
Earlier work this paper cites.
Observational modeling effects for movement dynamics and movement outcome measures across differing task constraints: A meta-analysis
Derek Ashford, Simon J. Bennett, and Keith Davids · 2006
Earlier work this paper cites.
Simbicon: Simple biped locomotion control
KangKang Yin, Kevin Loken, and Michiel Van de Panne · 2007
Earlier work this paper cites.
Sparse but not ‘grandmother-cell’ coding in the medial temporal lobe
R. Quian Quiroga, G. Kreiman, C. Koch, and I. Fried · 2007
Earlier work this paper cites.
Autonomous learning of stable quadruped locomotion
Manish Saggar, Thomas D’Silva, Nate Kohl, and Peter Stone · 2007
Earlier work this paper cites.
The chin pinch: A case study in skill learning on a legged robot
Peggy Fidelman and Peter Stone · 2007
Earlier work this paper cites.
Half field offense in RoboCup soccer: A multiagent reinforcement learning case study
Shivaram Kalyanakrishnan, Yaxin Liu, and Peter Stone · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2008
Earlier work this paper cites.
Model-based reinforcement learning in a complex domain
Shivaram Kalyanakrishnan, Peter Stone, and Yaxin Liu · 2008
Earlier work this paper cites.
Learning to dribble on a real robot by success and failure
Martin A. Riedmiller, Roland Hafner, Sascha Lange, and Martin Lauer · 2008
Earlier work this paper cites.
Reinforcement learning for robot soccer
Martin Riedmiller, Thomas Gabel, Roland Hafner, and Sascha Lange · 2009
Cited alongside, same era.
Generalized biped walking control
Stelian Coros, Philippe Beaudoin, and Michiel Van de Panne · 2010
Cited alongside, same era.
The world of Independent learners is not Markovian
Guillaume J. Laurent, Laëtitia Matignon, and Nadine Le Fort-Piat · 2010
Cited alongside, same era.
Learning complementary multiagent behaviors: A case study
Shivaram Kalyanakrishnan and Peter Stone · 2010
Cited alongside, same era.
Cognitive concepts in autonomous soccer playing robots
Martin Lauer, Roland Hafner, Sascha Lange, and Martin Riedmiller · 2010
Cited alongside, same era.
On progress in RoboCup: The simulation league showcase
Thomas Gabel and Martin Riedmiller · 2010
Cited alongside, same era.
OpenAI Five
OpenAI · 2018
Later among the works it cites.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2018
Later among the works it cites.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2018
Later among the works it cites.
DeepMimic: Example-guided deep reinforcement learning of physics-based character skills
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel van de Panne · 2018
Later among the works it cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Rémi Munos, Nicolas Heess, and Martin A. Riedmiller · 2018
Later among the works it cites.
Beyond expected goals
William Spearman · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning powerful kicks on the Aibo ERS-7: The quest for a striker
Matthew Hausknecht and Peter Stone · 2011
Cited alongside, same era.
On optimizing interdependent skills: A case study in simulated 3d humanoid robot soccer
Daniel Urieli, Patrick MacAlpine, Shivaram Kalyanakrishnan, Yinon Bentor, and Peter Stone · 2011
Cited alongside, same era.
Evolution of cooperation among mammalian carnivores and its relevance to hominin evolution
Jennifer E. Smith, Eli M. Swanson, Daphna Reed, and Kay E. Holekamp · 2012
Cited alongside, same era.
Terrain runner: control, parameterization, composition, and planning for highly dynamic motions
Libin Liu, KangKang Yin, Michiel van de Panne, and Baining Guo · 2012
Cited alongside, same era.
Developmental activities and the acquisition of superior anticipation and decision making in soccer players
André Roca, A Mark Williams, and Paul R Ford · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy P. Lillicrap, and Martin A. Riedmiller · 2018
Later among the works it cites.
Learning by playing - solving sparse reward tasks from scratch
Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom van de Wiele, Vlad Mnih, Nicolas Heess, and Jost Tobias Springenberg · 2018
Later among the works it cites.
Re-evaluating evaluation
David Balduzzi, Karl Tuyls, Julien Pérolat, and Thore Graepel · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Shixiang (Shane) Gu, Honglak Lee, and Sergey Levine · 2018
Later among the works it cites.
Learning basketball dribbling skills using trajectory optimization and deep reinforcement learning
Libin Liu and Jessica Hodgins · 2018
Later among the works it cites.
Physics-based motion capture imitation with deep reinforcement learning
Nuttapong Chentanez, Matthias Müller, Miles Macklin, Viktor Makoviychuk, and Stefan Jeschke · 2018
Later among the works it cites.
Counterfactual multi-agent policy gradients
Jakob N. Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Later among the works it cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yura Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Later among the works it cites.
Learning an embedding space for transferable robot skills
Karol Hausman, Jost Tobias Springenberg, Ziyu Wang, Nicolas Heess, and Martin Riedmiller · 2018
Later among the works it cites.
Latent space policies for hierarchical reinforcement learning
Tuomas Haarnoja, Kristian Hartikainen, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Meta learning shared hierarchies
Kevin Frans, Jonathan Ho, Xi Chen, Pieter Abbeel, and John Schulman · 2018
Later among the works it cites.
Overlapping layered learning
Patrick MacAlpine and Peter Stone · 2018
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M. Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castañeda, Charles Beattie, Neil C. Rabinowitz, Ari S. Morcos, Avraham Ruderman, Nicolas Sonnerat, Tim Green, Louise Deason, Joel Z. Leibo, David Silver, Demis Hassabis, Koray Kavukcuoglu, and Thore Graepel · 2019
Later among the works it cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Abhishek Gupta, Vikash Kumar, Corey Lynch, Sergey Levine, and Karol Hausman · 2019
Later among the works it cites.
Hierarchical visuomotor control of humanoids
Josh Merel, Arun Ahuja, Vu Pham, Saran Tunyasuvunakool, Siqi Liu, Dhruva Tirumala, Nicolas Heess, and Greg Wayne · 2019
Later among the works it cites.
MCP: learning composable hierarchical control with multiplicative compositional policies
Xue Bin Peng, Michael Chang, Grace Zhang, Pieter Abbeel, and Sergey Levine · 2019
Later among the works it cites.
Neural probabilistic motor primitives for humanoid control
Josh Merel, Leonard Hasenclever, Alexandre Galashov, Arun Ahuja, Vu Pham, Greg Wayne, Yee Whye Teh, and Nicolas Heess · 2019
Later among the works it cites.
Emergent coordination through competition
Siqi Liu, Guy Lever, Josh Merel, Saran Tunyasuvunakool, Nicolas Heess, and Thore Graepel · 2019
Later among the works it cites.
Information asymmetry in KL-regularized RL
Alexandre Galashov, Siddhant Jayakumar, Leonard Hasenclever, Dhruva Tirumala, Jonathan Schwarz, Guillaume Desjardins, Wojtek M. Czarnecki, Yee Whye Teh, Razvan Pascanu, and Nicolas Heess · 2019
Later among the works it cites.
Scalable muscle-actuated human simulation and control
Seunghwan Lee, Moonseok Park, Kyoungmin Lee, and Jehee Lee · 2019
Later among the works it cites.
Learning predict-and-simulate policies from unorganized human motion data
Soohwan Park, Hoseok Ryu, Seyoung Lee, Sunmin Lee, and Jehee Lee · 2019
Later among the works it cites.
Drecon: data-driven responsive control of physics-based characters
Kevin Bergamin, Simon Clavet, Daniel Holden, and James Richard Forbes · 2019
Later among the works it cites.
Learning to sit: Synthesizing human-chair interactions via hierarchical control
Yu-Wei Chao, Jimei Yang, Weifeng Chen, and Jia Deng · 2019
Later among the works it cites.
More Parkour Atlas, 2019
Boston Dynamics · 2019
Later among the works it cites.
Learning agile and dynamic motor skills for legged robots
Jemin Hwangbo, Joonho Lee, Alexey Dosovitskiy, Dario Bellicoso, Vassilios Tsounis, Vladlen Koltun, and Marco Hutter · 2019
Later among the works it cites.
A survey and critique of multiagent deep reinforcement learning
Pablo Hernandez-Leal, Bilal Kartal, and Matthew E. Taylor · 2019
Later among the works it cites.
Near-optimal representation learning for hierarchical reinforcement learning
Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine · 2019
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2019
Later among the works it cites.
B-Human 2019 – complex team play under natural lighting conditions
Thomas Röfer, Tim Laue, Gerrit Felsch, Arne Hasselbring, Tim Haß, Jan Oppermann, Philip Reichenberg, and Nicole Schrader · 2019
Later among the works it cites.
Learning to run faster in a humanoid robot soccer environment through reinforcement learning
Miguel Abreu, Luis Paulo Reis, and Nuno Lau · 2019
Later among the works it cites.
UT Austin Villa: RoboCup 2019 3D simulation league competition and technical challenge champions
Patrick MacAlpine, Faraz Torabi, Brahma Pavse, and Peter Stone · 2019
Later among the works it cites.
Open-ended learning in symmetric zero-sum games
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Pérolat, Max Jaderberg, and Thore Graepel · 2019
Later among the works it cites.
Catch & Carry: Reusable Neural Controllers for Vision-Guided Whole-Body Tasks
Josh Merel, Saran Tunyasuvunakool, Arun Ahuja, Yuval Tassa, Leonard Hasenclever, Vu Pham, Tom Erez, Greg Wayne, and Nicolas Heess · 2020
Later among the works it cites.
Learning quadrupedal locomotion over challenging terrain
Joonho Lee, Jemin Hwangbo, Lorenz Wellhausen, Vladlen Koltun, and Marco Hutter · 2020
Later among the works it cites.
Learning agile robotic locomotion skills by imitating animals, 2020
Xue Bin Peng, Erwin Coumans, Tingnan Zhang, Tsang-Wei Lee, Jie Tan, and Sergey Levine · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control
Yuval Tassa, Saran Tunyasuvunakool, Alistair Muldal, Yotam Doron, Siqi Liu, Steven Bohez, Josh Merel, Tom Erez, Timothy Lillicrap, and Nicolas Heess · 2020
Later among the works it cites.
Game plan: What AI can do for football, and what football can do for AI, 2020
Karl Tuyls, Shayegan Omidshafiei, Paul Muller, Zhe Wang, Jerome Connor, Daniel Hennes, Ian Graham, William Spearman, Tim Waskett, Dafydd Steele, Pauline Luc, Adria Recasens, Alexandre Galashov, Gregory Thornton, Romuald Elie, Pablo Sprechmann, Pol Moreno, Kris Cao, Marta Garnelo, Praneet Dutta, Michal Valko, Nicolas Heess, Alex Bridgland, Julien Perolat, Bart De Vylder, Ali Eslami, Mark Rowland, Andrew Jaegle, Remi Munos, Trevor Back, Razia Ahamed, Simon Bouton, Nathalie Beauguerlange, Jackson Broshear, Thore Graepel, and Demis Hassabis · 2020
Later among the works it cites.
CoMic: Complementary task learning & mimicry for reusable skills
Leonard Hasenclever, Fabio Pardo, Raia Hadsell, Nicolas Heess, and Josh Merel · 2020
Later among the works it cites.
Behavior priors for efficient reinforcement learning, 2020
Dhruva Tirumala, Alexandre Galashov, Hyeonwoo Noh, Leonard Hasenclever, Razvan Pascanu, Jonathan Schwarz, Guillaume Desjardins, Wojciech Marian Czarnecki, Arun Ahuja, Yee Whye Teh, and Nicolas Heess · 2020
Later among the works it cites.
Real world games look like spinning tops
Wojciech M. Czarnecki, Gauthier Gidel, Brendan Tracey, Karl Tuyls, Shayegan Omidshafiei, David Balduzzi, and Max Jaderberg · 2020
Later among the works it cites.
An overview of multi-agent reinforcement learning from game theoretical perspective, 2020
Yaodong Yang and Jun Wang · 2020
Later among the works it cites.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2020
Later among the works it cites.
Compositional transfer in hierarchical reinforcement learning
Markus Wulfmeier, Abbas Abdolmaleki, Roland Hafner, Jost Tobias Springenberg, Michael Neunert, Noah Siegel, Tim Hertweck, Thomas Lampe, Nicolas Heess, and Martin Riedmiller · 2020
Later among the works it cites.
Reinforced grounded action transformation for sim-to-real transfer
Haresh Karnan, Siddharth Desai, Josiah P. Hanna, Garrett Warnell, and Peter Stone · 2020
Later among the works it cites.
Stochastic grounded action transformation for robot learning in simulation
Siddharth Desai, Haresh Karnan, Josiah P. Hanna, Garrett Warnell, and Peter Stone · 2020
Later among the works it cites.