Fetching the paper…
Reading the bibliography…
This paper investigates a population-based training regime based on game-theoretic principles called Policy-Spaced Response Oracles (PSRO).
A simplified two-person poker
Harold W Kuhn · 1950
Earlier work this paper cites.
Equilibrium points in n n -person games
John F Nash · 1950
Earlier work this paper cites.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
An iterative method of solving a game
Julia Robinson · 1951
Earlier work this paper cites.
The logic of animal conflicts
J. Maynard Smith and G. R. Price · 1973
Earlier work this paper cites.
The rating of chessplayers, past and present
Arpad E Elo · 1978
Earlier work this paper cites.
Evolutionary stable strategies and game dynamics
Peter D Taylor and Leo B Jonker · 1978
Earlier work this paper cites.
Replicator dynamics
Peter Schuster and Karl Sigmund · 1983
Earlier work this paper cites.
A general theory of equilibrium selection in games
John C Harsanyi, Reinhard Selten, et al · 1988
Earlier work this paper cites.
Analyzing complex strategic interactions in multi-agent systems
William E Walsh, Rajarshi Das, Gerald Tesauro, and Jeffrey O Kephart · 2002
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H. Brendan McMahan, Geoffrey J. Gordon, and Avrim Blum · 2003
Earlier work this paper cites.
Choosing samples to compute heuristic-strategy Nash equilibrium
William E Walsh, David C Parkes, and Rajarshi Das · 2003
Earlier work this paper cites.
An evolutionary game-theoretic comparison of two double-auction market designs
Steve Phelps, Simon Parsons, and Peter McBurney · 2004
Earlier work this paper cites.
Computing best-response strategies in infinite games of incomplete information
Daniel M Reeves and Michael P Wellman · 2004
Earlier work this paper cites.
Bayes’ bluff: Opponent modelling in poker
Finnegan Southey, Michael Bowling, Bryce Larson, Carmelo Piccione, Neil Burch, Darse Billings, and Chris Rayner · 2005
Earlier work this paper cites.
The parallel Nash memory for asymmetric games
Frans A. Oliehoek, Edwin D. de Jong, and Nikos Vlassis · 2006
Earlier work this paper cites.
Methods for empirical game-theoretic analysis
Michael P Wellman · 2006
Earlier work this paper cites.
What evolutionary game theory tells us about multiagent learning
Karl Tuyls and Simon Parsons · 2007
Earlier work this paper cites.
The complexity of computing a Nash equilibrium
Constantinos Daskalakis, Paul W Goldberg, and Christos H Papadimitriou · 2009
Cited alongside, same era.
Probabilistic analysis of simulation-based games
Yevgeniy Vorobeychik · 2010
Cited alongside, same era.
Independent reinforcement learners in cooperative Markov games: A survey regarding coordination problems
Laetitia Matignon, Guillaume J. Laurent, and Nadine Le Fort-Piat · 2012
Cited alongside, same era.
Projected dynamical systems and variational inequalities with applications , volume 2
Anna Nagurney and Ding Zhang · 2012
Cited alongside, same era.
Scaling simulation-based game analysis through deviation-preserving reduction
Bryce Wiedenbeck and Michael P. Wellman · 2012
Cited alongside, same era.
On the complexity of approximating a Nash equilibrium
Constantinos Daskalakis · 2013
Peng Peng, Ying Wen, Yaodong Yang, Quan Yuan, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Later among the works it cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Later among the works it cites.
Counterfactual multi-agent policy gradients
Jakob N Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Later among the works it cites.
Lenient multi-agent deep reinforcement learning
Gregory Palmer, Karl Tuyls, Daan Bloembergen, and Rahul Savani · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The complexity of the homotopy method, equilibrium selection, and Lemke-Howson solutions
Paul W Goldberg, Christos H Papadimitriou, and Rahul Savani · 2013
Cited alongside, same era.
Evolutionary advantage of reciprocity in collision avoidance
Daniel Hennes, Daniel Claes, and Karl Tuyls · 2013
Cited alongside, same era.
The replicator equation and other game dynamics
Ross Cressman and Yi Tao · 2014
Cited alongside, same era.
Bootstrap statistics for empirical games
Bryce Wiedenbeck, Ben-Alexander Cassell, and Michael P. Wellman · 2014
Cited alongside, same era.
Evolutionary dynamics of multi-agent learning: A survey
Daan Bloembergen, Karl Tuyls, Daniel Hennes, and Michael Kaisers · 2015
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
A generalised method for empirical game theoretic analysis
Karl Tuyls, Julien Perolat, Marc Lanctot, Joel Z Leibo, and Thore Graepel · 2018
Later among the works it cites.
Multiagent soft q-learning
Ermo Wei, Drew Wicke, David Freelan, and Sean Luke · 2018
Later among the works it cites.
Open-ended learning in symmetric zero-sum games
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Pérolat, Max Jaderberg, and Thore Graepel · 2019
Closest in time.
A survey and critique of multiagent deep reinforcement learning
Pablo Hernandez-Leal, Bilal Kartal, and Matthew E Taylor · 2019
Closest in time.
Actor-attention-critic for multi-agent reinforcement learning
Shariq Iqbal and Fei Sha · 2019
Closest in time.
Human-level performance in 3D multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M. Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castañeda, Charles Beattie, Neil C. Rabinowitz, Ari S. Morcos, Avraham Ruderman, Nicolas Sonnerat, Tim Green, Louise Deason, Joel Z. Leibo, David Silver, Demis Hassabis, Koray Kavukcuoglu, and Thore Graepel · 2019
Closest in time.
Evolutionary reinforcement learning for sample-efficient multiagent coordination
Shauharda Khadka, Somdeb Majumdar, and Kagan Tumer · 2019
Closest in time.
OpenSpiel: A framework for reinforcement learning in games
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Pérolat, Sriram Srinivasan, Finbarr Timbers, Karl Tuyls, Shayegan Omidshafiei, Daniel Hennes, Dustin Morrill, Paul Muller, Timo Ewalds, Ryan Faulkner, János Kramár, Bart De Vylder, Brennan Saeta, James Bradbury, David Ding, Sebastian Borgeaud, Matthew Lai, Julian Schrittwieser, Thomas Anthony, Edward Hughes, Ivo Danihelka, and Jonah Ryan-Davis · 2019
Closest in time.
Emergent coordination through competition
Siqi Liu, Guy Lever, Josh Merel, Saran Tunyasuvunakool, Nicolas Heess, and Thore Graepel · 2019
Closest in time.
Beyond local nash equilibria for adversarial networks
Frans A. Oliehoek, Rahul Savani, Jose Gallego, Elise van der Pol, and Roderich Groß · 2019
Closest in time.
α \alpha -Rank: Multi-agent evaluation by evolution
Shayegan Omidshafiei, Christos Papadimitriou, Georgios Piliouras, Karl Tuyls, Mark Rowland, Jean-Baptiste Lespiau, Wojciech M Czarnecki, Marc Lanctot, Julien Perolat, and Rémi Munos · 2019
Closest in time.
Multiagent evaluation under incomplete information
Mark Rowland, Shayegan Omidshafiei, Karl Tuyls, Julien Perolat, Michal Valko, Georgios Piliouras, and Remi Munos · 2019
Closest in time.