Fetching the paper…
Reading the bibliography…
Policy Space Response Oracle methods (PSRO) provide a general solution to learn Nash equilibrium in two-player zero-sum games but suffer from two drawbacks: (1) the computation inefficiency due to the need for consistent meta-game evaluation via simulations, and (2) the exploration inefficiency due to finding the best response against a fixed meta-strategy at every epoch.
Game theory
Drew Fudenberg and Jean Tirole · 1991
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Yoav Freund and Robert E Schapire · 1999
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H Brendan McMahan, Geoffrey J Gordon, and Avrim Blum · 2003
Earlier work this paper cites.
Convergence and no-regret in multiagent learning
Michael Bowling · 2004
Earlier work this paper cites.
The cardinality matrix constraint
Jean-Charles Régin and Carla P Gomes · 2004
Earlier work this paper cites.
Mixed-integer programming methods for finding nash equilibria
Tuomas Sandholm, Andrew Gilpin, and Vincent Conitzer · 2005
Earlier work this paper cites.
Generalised weakened fictitious play
David S Leslie and Edmund J Collins · 2006
Earlier work this paper cites.
A new algorithm for generating equilibria in massive zero-sum games
Martin Zinkevich, Michael Bowling, and Neil Burch · 2007
Earlier work this paper cites.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2007
Earlier work this paper cites.
On range of skill
Thomas Dueholm Hansen, Peter Bro Miltersen, and Troels Bjerre Sørensen · 2008
Earlier work this paper cites.
Multiagent systems: algorithmic, game-theoretic, and logical foundations by y. shoham and k. leyton-brown cambridge university press, 2008
Haris Aziz · 2010
Earlier work this paper cites.
Near-optimal no-regret algorithms for zero-sum games
Constantinos Daskalakis, Alan Deckelbaum, and Anthony Kim · 2011
Earlier work this paper cites.
Accelerating best response calculation in large extensive games
Michael Johanson, Kevin Waugh, Michael Bowling, and Martin Zinkevich · 2011
Earlier work this paper cites.
Heads-up limit hold’em poker is solved
Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin · 2015
Earlier work this paper cites.
Strategy-based warm starting for regret minimization in games
Noam Brown and Tuomas Sandholm · 2016
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrūnas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Mean field multi-agent reinforcement learning
Yaodong Yang, Rui Luo, Minne Li, Ming Zhou, Weinan Zhang, and Jun Wang · 2018
Iterative empirical game solving via single policy best response
Max Smith, Thomas Anthony, and Michael Wellman · 2020
Later among the works it cites.
A deterministic linear program solver in current matrix multiplication time
Jan van den Brand · 2020
Later among the works it cites.
α \alpha α \alpha -rank: Practically scaling α \alpha -rank through stochastic optimisation
Yaodong Yang, Rasul Tutunov, Phu Sakulwongtana, and Haitham Bou Ammar · 2020
Later among the works it cites.
Q-mixing network for multi-agent pathfinding in partially observable grid environments
Vasilii Davydov, Alexey Skrynnik, Konstantin Yakovlev, and Aleksandr Panov · 2021
Later among the works it cites.
Online double oracle
Yaodong Yang Le Cong Dinh, Zheng Tian, Nicolas Perez Nieves, Oliver Slumbers, David Henry Mguni, Haitham Bou Ammar, and Jun Wang · 2021
Later among the works it cites.
Neupl: Neural population learning
Siqi Liu, Luke Marris, Daniel Hennes, Josh Merel, Nicolas Heess, and Thore Graepel · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Open-ended learning in symmetric zero-sum games
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Perolat, Max Jaderberg, and Thore Graepel · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
Distilling policy distillation
Wojciech M Czarnecki, Razvan Pascanu, Simon Osindero, Siddhant Jayakumar, Grzegorz Swirszcz, and Max Jaderberg · 2019
Cited alongside, same era.
A generalized framework for self-play training
Daniel Hernandez, Kevin Denamganaï, Yuan Gao, Peter York, Sam Devlin, Spyridon Samothrakis, and James Alfred Walker · 2019
Cited alongside, same era.
α \alpha -rank: Multi-agent evaluation by evolution
Shayegan Omidshafiei, Christos Papadimitriou, Georgios Piliouras, Karl Tuyls, Mark Rowland, Jean-Baptiste Lespiau, Wojciech M Czarnecki, Marc Lanctot, Julien Perolat, and Remi Munos · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Later among the works it cites.
Towards unifying behavioral and response diversity for open-ended learning in zero-sum games
Xiangyu Liu, Hangtian Jia, Ying Wen, Yaodong Yang, Yujing Hu, Yingfeng Chen, Changjie Fan, and Zhipeng Hu · 2021
Later among the works it cites.
Xdo: A double oracle algorithm for extensive-form games
Stephen McAleer, John Banister Lanier, Kevin A Wang, Pierre Baldi, and Roy Fox · 2021
Later among the works it cites.
Modelling behavioural diversity for learning in open-ended games
Nicolas Perez-Nieves, Yaodong Yang, Oliver Slumbers, David H Mguni, Ying Wen, and Jun Wang · 2021
Later among the works it cites.
Evaluating strategy exploration in empirical game-theoretic analysis
Yongzhao Wang, Qiurui Ma, and Michael P Wellman · 2021
Later among the works it cites.
Diverse auto-curriculum is critical for successful real-world multiagent learning systems
Yaodong Yang, Jun Luo, Ying Wen, Oliver Slumbers, Daniel Graves, Haitham Bou Ammar, Jun Wang, and Matthew E Taylor · 2021
Later among the works it cites.
Shicong Cen, Fan Chen, and Yuejie Chi · 2022
Closest in time.
Anytime optimal psro for two-player zero-sum games, 2022
Stephen McAleer, Kevin Wang, Marc Lanctot, John Lanier, Pierre Baldi, and Roy Fox · 2022
Closest in time.