Fetching the paper…
Reading the bibliography…
Open-ended learning methods that automatically generate a curriculum of increasingly challenging tasks serve as a promising avenue toward generally capable reinforcement learning agents.
Emergent tool use from multi-agent autocurricula, 2019
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 1909
Earlier work this paper cites.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
Stochastic games
L. S. Shapley · 1953
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Evolutionary robotics and the radical envelope-of-noise hypothesis
Nick Jakobi · 1997
Earlier work this paper cites.
Generalised weakened fictitious play
David S Leslie and Edmund J Collins · 2006
Earlier work this paper cites.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2007
Earlier work this paper cites.
Griddly: A platform for ai research in games, 2020
Chris Bamford, Shengyi Huang, and Simon Lucas · 2011
Earlier work this paper cites.
Fictitious self-play in extensive-form games
Johannes Heinrich, Marc Lanctot, and David Silver · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
Transferring end-to-end visuomotor control from simulation to real world for a multi-stage task
Stephen James, Andrew J. Davison, and Edward Johns · 2017
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning, 2017
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Perolat, David Silver, and Thore Graepel · 2017
Earlier work this paper cites.
Multi-agent reinforcement learning in sequential social dilemmas, 2017
Joel Z. Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel · 2017
Earlier work this paper cites.
CAD2RL: real single-image flight without a single real image
Fereshteh Sadeghi and Sergey Levine · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Open-endedness: The last grand challenge you’ve never heard of
Kenneth O Stanley, Joel Lehman, and Lisa Soros · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Cited alongside, same era.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2018
Cited alongside, same era.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Cited alongside, same era.
Open-ended learning in symmetric zero-sum games
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Perolat, Max Jaderberg, and Thore Graepel · 2019
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver · 2020
Later among the works it cites.
Enhanced POET: Open-ended reinforcement learning through unbounded invention of learning challenges and their solutions
Rui Wang, Joel Lehman, Aditya Rawal, Jiale Zhi, Yulun Li, Jeffrey Clune, and Kenneth Stanley · 2020
Later among the works it cites.
Self-paced context evaluation for contextual reinforcement learning
Theresa Eimer, André Biedenkapp, Frank Hutter, and Marius Lindauer · 2021
Later among the works it cites.
Neural Auto-Curricula in Two-Player Zero-Sum Games
Xidong Feng, Oliver Slumbers, Ziyu Wan, Bo Liu, Stephen McAleer, Ying Wen, Jun Wang, and Yaodong Yang · 2021
Later among the works it cites.
Pick your battles: Interaction graphs as population-level objectives for strategic diversity, 2021
Marta Garnelo, Wojciech Marian Czarnecki, Siqi Liu, Dhruva Tirumala, Junhyuk Oh, Gauthier Gidel, Hado van Hasselt, and David Balduzzi · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, Rafal Józefowicz, Scott Gray, Catherine Olsson, Jakub Pachocki, Michael Petrov, Henrique Pondé de Oliveira Pinto, Jonathan Raiman, Tim Salimans, Jeremy Schlatter, Jonas Schneider, Szymon Sidor, Ilya Sutskever, Jie Tang, Filip Wolski, and Susan Zhang · 2019
Cited alongside, same era.
Superhuman ai for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Cited alongside, same era.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M. Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castañ eda, Charles Beattie, Neil C. Rabinowitz, Ari S. Morcos, Avraham Ruderman, Nicolas Sonnerat, Tim Green, Louise Deason, Joel Z. Leibo, David Silver, Demis Hassabis, Koray Kavukcuoglu, and Thore Graepel · 2019
Cited alongside, same era.
Self-paced contextual reinforcement learning
Pascal Klink, Hany Abdulsamad, Boris Belousov, and Jan Peters · 2019
Cited alongside, same era.
Joel Z. Leibo, Edward Hughes, Marc Lanctot, and Thore Graepel · 2019
Cited alongside, same era.
Car racing with pytorch
Xiaoteng Ma · 2019
Cited alongside, same era.
Teacher algorithms for curriculum learning of deep RL in continuously parameterized environments
Rémy Portelas, Cédric Colas, Katja Hofmann, and Pierre-Yves Oudeyer · 2019
Cited alongside, same era.
Later among the works it cites.
Environment generation for zero-shot compositional reinforcement learning
Izzeddin Gur, Natasha Jaques, Yingjie Miao, Jongwook Choi, Manoj Tiwari, Honglak Lee, and Aleksandra Faust · 2021
Later among the works it cites.
Exploration-exploitation in multi-agent competition: Convergence with bounded rationality
Stefanos Leonardos, Georgios Piliouras, and Kelly Spendlove · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
Open Ended Learning Team, Adam Stooke, Anuj Mahajan, Catarina Barros, Charlie Deck, Jakob Bauer, Jakub Sygnowski, Maja Trebacz, Max Jaderberg, Michaël Mathieu, Nat McAleese, Nathalie Bradley-Schmieg, Nathaniel Wong, Nicolas Porcel, Roberta Raileanu, Steph Hughes-Fitt, Valentin Dalibard, and Wojciech Marian Czarnecki · 2021
Later among the works it cites.
Deep latent competition: Learning to race using visual control policies in latent space, 2021
Wilko Schwarting, Tim Seyde, Igor Gilitschenski, Lucas Liebenwein, Ryan Sander, Sertac Karaman, and Daniela Rus · 2021
Later among the works it cites.
Diverse auto-curriculum is critical for successful real-world multiagent learning systems
Yaodong Yang, Jun Luo, Ying Wen, Oliver Slumbers, Daniel Graves, Haitham Bou-Ammar, Jun Wang, and Matthew E. Taylor · 2021
Later among the works it cites.
GriddlyJS: A web IDE for reinforcement learning
Christopher Bamford, Minqi Jiang, Mikayel Samvelyan, and Tim Rocktäschel · 2022
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan D. Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de Las Casas, Craig Donner, Leslie Fritz, Cristian Galperti, Andrea Huber, James Keeling, Maria Tsimpoukelli, Jackie Kay, Antoine Merle, J-M. Moret, Seb Noury, Federico Pesamosca, David G. Pfau, Olivier Sauter, Cristian Sommariva, Stefano Coda, B. Duval, Ambrogio Fasoli, Pushmeet Kohli, Koray Kavukcuoglu, Demis Hassabis, and Martin A. Riedmiller · 2022
Later among the works it cites.
SMACv2: An improved benchmark for cooperative multi-agent reinforcement learning, 2022
Benjamin Ellis, Skander Moalla, Mikayel Samvelyan, Mingfei Sun, Anuj Mahajan, Jakob N. Foerster, and Shimon Whiteson · 2022
Later among the works it cites.
Generalization in cooperative multi-agent systems
Anuj Mahajan, Mikayel Samvelyan, Tarun Gupta, Benjamin Ellis, Mingfei Sun, Tim Rocktäschel, and Shimon Whiteson · 2022
Later among the works it cites.
Evolving curricula with regret-based environment design, 2022
Jack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan, Jakob Foerster, Edward Grefenstette, and Tim Rocktäschel · 2022
Later among the works it cites.
Outracing champion Gran Turismo drivers with deep reinforcement learning
Peter R. Wurman, Samuel Barrett, Kenta Kawamoto, James MacGlashan, Kaushik Subramanian, Thomas J. Walsh, Roberto Capobianco, Alisa Devlic, Franziska Eckert, Florian Fuchs, Leilani Gilpin, Piyush Khandelwal, Varun Kompella, HaoChih Lin, Patrick MacAlpine, Declan Oller, Takuma Seno, Craig Sherstan, Michael D. Thomure, Houmehr Aghabozorgi, Leon Barrett, Rory Douglas, Dion Whitehead, Peter Dürr, Peter Stone, Michael Spranger, and Hiroaki Kitano · 2022
Later among the works it cites.