Fetching the paper…
Reading the bibliography…
Self-play (SP) is a popular multi-agent reinforcement learning (MARL) framework for solving competitive games, where each agent optimizes policy by treating others as part of the environment.
Iterative solution of games by fictitious play
George W Brown. 1951 · 1951
Earlier work this paper cites.
The rating of chessplayers, past and present
Arpad E Elo. 1978 · 1978
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning. In Proceedings of the eleventh international conference on machine learning , Vol. 157. 157–163
Michael L Littman. 1994 · 1994
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Gerald Tesauro. 1994 · 1994
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Yoav Freund and Robert E Schapire. 1999 · 1999
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary. In Proceedings of the 20th International Conference on Machine Learning (ICML-03) . 536–543
H Brendan McMahan, Geoffrey J Gordon, and Avrim Blum. 2003 · 2003
Earlier work this paper cites.
Learning, regret minimization, and equilibria
Avrim Blum and Yishay Monsour. 2007 · 2007
Earlier work this paper cites.
Online learning and online convex optimization
Shai Shalev-Shwartz et al · 2012
Earlier work this paper cites.
Characterization and computation of local Nash equilibria in continuous games. In 2013 51st Annual Allerton Conference on Communication, Control, and Computing (Allerton) . IEEE, 917–924
Lillian J Ratliff, Samuel A Burden, and S Shankar Sastry. 2013 · 2013
Earlier work this paper cites.
Fictitious self-play in extensive-form games. In International conference on machine learning . PMLR, 805–813
Johannes Heinrich, Marc Lanctot, and David Silver. 2015 · 2015
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver. 2016 · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel. 2017 · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning. In International conference on machine learning . PMLR, 4295–4304
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
The hanabi challenge: A new frontier for ai research
Nolan Bard, Jakob N Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, et al · 2020
Later among the works it cites.
Neural replicator dynamics: Multiagent learning via hedging policy gradients. In Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems . 492–501
Daniel Hennes, Dustin Morrill, Shayegan Omidshafiei, Rémi Munos, Julien Perolat, Marc Lanctot, Audrunas Gruslys, Jean-Baptiste Lespiau, Paavo Parmas, Edgar Duéñez-Guzmán, et al · 2020
Later among the works it cites.
Google research football: A novel reinforcement learning environment. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 4501–4510
Karol Kurach, Anton Raichuk, Piotr Stańczyk, Michał Zając, Olivier Bachem, Lasse Espeholt, Carlos Riquelme, Damien Vincent, Marcin Michalski, Olivier Bousquet, et al · 2020
Later among the works it cites.
Le Cong Dinh, Yaodong Yang, Zheng Tian, Nicolas Perez Nieves, Oliver Slumbers, David Henry Mguni, Haitham Bou Ammar, and Jun Wang. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lianmin Zheng, Jiacheng Yang, Han Cai, Ming Zhou, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Cited alongside, same era.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch. 2019 · 2019
Cited alongside, same era.
Open-ended learning in symmetric zero-sum games. In International Conference on Machine Learning . PMLR, 434–443
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Perolat, Max Jaderberg, and Thore Graepel. 2019 · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
Deep counterfactual regret minimization. In International conference on machine learning . PMLR, 793–802
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm. 2019 · 2019
Cited alongside, same era.
Human-level performance in 3D multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castaneda, Charles Beattie, Neil C Rabinowitz, Ari S Morcos, Avraham Ruderman, et al · 2019
Cited alongside, same era.
A generalized training approach for multiagent learning
Paul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls, Julien Perolat, Siqi Liu, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes, et al · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
TiKick: Towards Playing Multi-agent Football Full Games from Single-agent Demonstrations
Shiyu Huang, Wenze Chen, Longfei Zhang, Ziyang Li, Fengming Zhu, Deheng Ye, Ting Chen, and Jun Zhu. 2021 · 2021
Later among the works it cites.
Towards Unifying Behavioral and Response Diversity for Open-ended Learning in Zero-sum Games
Xiangyu Liu, Hangtian Jia, Ying Wen, Yujing Hu, Yingfeng Chen, Changjie Fan, Zhipeng Hu, and Yaodong Yang. 2021 · 2021
Later among the works it cites.
XDO: A double oracle algorithm for extensive-form games
Stephen McAleer, John B Lanier, Kevin A Wang, Pierre Baldi, and Roy Fox. 2021 · 2021
Later among the works it cites.
Modelling behavioural diversity for learning in open-ended games. In International Conference on Machine Learning . PMLR, 8514–8524
Nicolas Perez-Nieves, Yaodong Yang, Oliver Slumbers, David H Mguni, Ying Wen, and Jun Wang. 2021 · 2021
Later among the works it cites.
From Poincaré recurrence to convergence in imperfect information games: Finding equilibrium via regularization. In International Conference on Machine Learning . PMLR, 8525–8535
Julien Perolat, Remi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland, Pedro Ortega, Neil Burch, Thomas Anthony, David Balduzzi, Bart De Vylder, et al · 2021
Later among the works it cites.
The surprising effectiveness of ppo in cooperative, multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre Bayen, and Yi Wu. 2021 · 2021
Later among the works it cites.
Anytime PSRO for Two-Player Zero-Sum Games
Stephen McAleer, Kevin Wang, JB Lanier, Marc Lanctot, Pierre Baldi, Tuomas Sandholm, and Roy Fox. 2022 · 2022
Later among the works it cites.
Samuel Sokota, Ryan D’Orazio, J Zico Kolter, Nicolas Loizou, Marc Lanctot, Ioannis Mitliagkas, Noam Brown, and Christian Kroer. 2022 · 2022
Later among the works it cites.