Fetching the paper…
Reading the bibliography…
Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a self-supervised autocurriculum inducing multiple distinct rounds of emergent strategy, many of which require sophisticated tool use and coordination.
Joel Z Leibo, Edward Hughes, Marc Lanctot, and Thore Graepel · 1903
Earlier work this paper cites.
Relational deep reinforcement learning
Vinicius Zambaldi, David Raposo, Adam Santoro, Victor Bapst, Yujia Li, Igor Babuschkin, Karl Tuyls, David Reichert, Timothy Lillicrap, Edward Lockhart, et al · 1909
Earlier work this paper cites.
The rating of chessplayers, past and present
A.E. Elo · 1978
Earlier work this paper cites.
Arms races between and within species
Richard Dawkins and John Richard Krebs · 1979
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Evolution, ecology and optimization of digital organisms
Thomas S Ray · 1992
Earlier work this paper cites.
Computational genetics, physiology, metabolism, neural systems, learning, vision, and behavior or poly world: Life in a new context
Larry Yaeger · 1994
Earlier work this paper cites.
Coevolutionary computation
Jan Paredis · 1995
Earlier work this paper cites.
Methods for competitive co-evolution: Finding opponents worth beating
Christopher D Rosin and Richard K Belew · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Manufacture and use of hook-tools by new caledonian crows
Gavin R Hunt · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Coevolution of a backgammon player
Jordan B Pollack, Alan D Blair, and Mark Land · 1997
Earlier work this paper cites.
Perpetuating evolutionary emergence
AD Channon, RI Damper, et al · 1998
Earlier work this paper cites.
Relational reinforcement learning
Sašo Džeroski, Luc De Raedt, and Kurt Driessens · 2001
Earlier work this paper cites.
Avida: A software platform for research in computational evolutionary biology
Charles Ofria and Claus O Wilke · 2004
Earlier work this paper cites.
Competitive coevolution through evolutionary complexification
Kenneth O Stanley and Risto Miikkulainen · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G Barto, and Satinder P Singh · 2005
Earlier work this paper cites.
Trueskill™: a bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel · 2007
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Carlos Diuk, Andre Cohen, and Michael L Littman · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Intrinsically motivated reinforcement learning: An evolutionary perspective
Satinder Singh, Richard L Lewis, Andrew G Barto, and Jonathan Sorg · 2010
Earlier work this paper cites.
Animal tool behavior: the use and manufacture of tools by animals
Robert W Shumaker, Kristina R Walkup, and Benjamin B Beck · 2011
Earlier work this paper cites.
Core cognition and beyond: The acquisition of physical and numerical knowledge
Renée Baillargeon and Susan Carey · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Identifying necessary conditions for open-ended evolution through the artificial life world of chromaria
L Soros and Kenneth Stanley · 2014
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2015
Cited alongside, same era.
Asymmetric actor critic for image-based robot learning
Lerrel Pinto, Marcin Andrychowicz, Peter Welinder, Wojciech Zaremba, and Pieter Abbeel · 2017
Later among the works it cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Requirements for open-ended evolution in natural and artificial systems
Tim Taylor · 2015
Cited alongside, same era.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Learning dexterous in-hand manipulation
Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2018
Later among the works it cites.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2018
Later among the works it cites.
Counterfactual multi-agent policy gradients
Jakob N Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Later among the works it cites.
Learning to play with intrinsically-motivated, self-aware agents
Nick Haber, Damian Mrowca, Stephanie Wang, Li F Fei-Fei, and Daniel L Yamins · 2018
Later among the works it cites.
Joel Lehman, Jeff Clune, Dusan Misevic, Christoph Adami, Lee Altenberg, Julie Beaulieu, Peter J Bentley, Samuel Bernard, Guillaume Beslon, David M Bryson, et al · 2018
Later among the works it cites.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2018
Later among the works it cites.
OpenAI Five
OpenAI · 2018
Later among the works it cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sainbayar Sukhbaatar, Zeming Lin, Ilya Kostrikov, Gabriel Synnaeve, Arthur Szlam, and Rob Fergus · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
The tools challenge: Rapid trial-and-error learning in physical problem solving
Kelsey R Allen, Kevin A Smith, and Joshua B Tenenbaum · 2019
Closest in time.
Structured agents for physical construction
Victor Bapst, Alvaro Sanchez-Gonzalez, Carl Doersch, Kimberly Stachenfeld, Pushmeet Kohli, Peter Battaglia, and Jessica Hamrick · 2019
Closest in time.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M. Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castañeda, Charles Beattie, Neil C. Rabinowitz, Ari S. Morcos, Avraham Ruderman, Nicolas Sonnerat, Tim Green, Louise Deason, Joel Z. Leibo, David Silver, Demis Hassabis, Koray Kavukcuoglu, and Thore Graepel · 2019
Closest in time.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro Ortega, Dj Strouse, Joel Z Leibo, and Nando De Freitas · 2019
Closest in time.
Emergent coordination through competition
Siqi Liu, Guy Lever, Nicholas Heess, Josh Merel, Saran Tunyasuvunakool, and Thore Graepel · 2019
Closest in time.
AlphaStar: Mastering the real-time strategy game StarCraft II
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Max Jaderberg, et al · 2019
Closest in time.
Rui Wang, Joel Lehman, Jeff Clune, and Kenneth O Stanley · 2019
Closest in time.
Improvisation through physical understanding: Using novel objects as tools with visual foresight
Annie Xie, Frederik Ebert, Sergey Levine, and Chelsea Finn · 2019
Closest in time.