Fetching the paper…
Reading the bibliography…
Recent advances in deep reinforcement learning (RL) have led to considerable progress in many 2-player zero-sum games, such as Go, Poker and Starcraft.
The game theory of play and integral equations with skew symmetric kernels (la théorie du jeu et les équations intégrales à noyau symétrique)
É. Borel · 1921
Earlier work this paper cites.
Probable inference, the law of succession, and statistical inference
Edwin B Wilson · 1927
Earlier work this paper cites.
Zur Theorie der Gesellschaftsspiele
J Von Neumann · 1928
Earlier work this paper cites.
Theory of Games and Economic Behavior (Commemorative Edition)
John von Neumann, Oskar Morgenstern, and Harold William Kuhn · 1944
Earlier work this paper cites.
Equilibrium points in n-person games
John F Nash et al · 1950
Earlier work this paper cites.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
An iterative method of solving a game
Julia Robinson · 1951
Earlier work this paper cites.
Note on a computation method in the theory of games
H.N. Shapiro · 1958
Earlier work this paper cites.
Diplomacy. Board Game
AB Calhamer · 1959
Earlier work this paper cites.
Mathematical Methods and Theory in Games, Programming and Economics
Samuel Karlin · 1959
Earlier work this paper cites.
Some studies in machine learning using the game of Checkers
Arthur L Samuel · 1959
Earlier work this paper cites.
The Greenblatt Chess Program
Richard D Greenblatt, Donald E Eastlake III, and Stephen D Crocker · 1967
Earlier work this paper cites.
An analysis of Alpha-Beta pruning
Donald E Knuth and Ronald W Moore · 1975
Earlier work this paper cites.
Studies in machine cognition using the game of Poker
Nicholas V Findler · 1977
Earlier work this paper cites.
Strategically zero-sum games: the class of games whose completely mixed equilibria cannot be improved upon
Hervé Moulin and J-P Vial · 1978
Earlier work this paper cites.
Evolutionary stable strategies and game dynamics
Peter D Taylor and Leo B Jonker · 1978
Earlier work this paper cites.
Backgammon computer program beats world champion
Hans J Berliner · 1980
Earlier work this paper cites.
The evolution of cooperation
Robert Axelrod and William Donald Hamilton · 1981
Earlier work this paper cites.
A world-championship-level Othello program
Paul S Rosenbloom · 1982
Earlier work this paper cites.
Evolution and the Theory of Games
John Maynard Smith · 1982
Earlier work this paper cites.
Diplomat, an agent in a multi agent environment: An overview
Sarit Kraus and Daniel Lehmann · 1988
Earlier work this paper cites.
An automated Diplomacy player
Sarit Kraus, Daniel Lehmann, and Eithan Ephrati · 1989
Earlier work this paper cites.
Thoughts on programming a diplomat
Michael R Hall and Daniel E Loeb · 1992
Earlier work this paper cites.
A world championship caliber Checkers program
Jonathan Schaeffer, Joseph Culberson, Norman Treloar, Brent Knight, Paul Lu, and Duane Szafron · 1992
Earlier work this paper cites.
Learning mixed equilibria
D. Fudenberg and D. M. Kreps · 1993
Earlier work this paper cites.
Three problems in learning mixed-strategy Nash equilibria
J. S. Jordan · 1993
Earlier work this paper cites.
Negotiation in a non-cooperative environment
Sarit Kraus, Eithan Ephrati, and Daniel Lehmann · 1994
Earlier work this paper cites.
On the complexity of the parity argument and other inefficient proofs of existence
Christos H Papadimitriou · 1994
Earlier work this paper cites.
Temporal difference learning of position evaluation in the game of Go
Nicol N Schraudolph, Peter Dayan, and Terrence J Sejnowski · 1994
Earlier work this paper cites.
TD-Gammon, a self-teaching Backgammon program, achieves master-level play
Gerald Tesauro · 1994
Earlier work this paper cites.
Confidence intervals for weighted proportions
Jennifer L Waller, Cheryl L Addy, Kirby L Jackson, and Carol Z Garrison · 1994
Earlier work this paper cites.
Designing and building a negotiating automated agent
Sarit Kraus and Daniel Lehmann · 1995
Earlier work this paper cites.
Quantal response equilibria for normal form games
Richard D McKelvey and Thomas R Palfrey · 1995
Earlier work this paper cites.
Polynomial-complexity deadlock avoidance policies for sequential resource allocation systems
Spiridon A Reveliotis, Mark A Lawley, and Placid M Ferreira · 1997
Earlier work this paper cites.
The Theory of Learning in Games
Drew Fudenberg, Fudenberg Drew, David K Levine, and David K Levine · 1998
Earlier work this paper cites.
On the rate of convergence of continuous-time fictitious play
Christopher Harris · 1998
Earlier work this paper cites.
A weakened form of fictitious play in two-person zero-sum games
Ben Van der Genugten · 2000
Earlier work this paper cites.
Deadlock avoidance in sequential resource allocation systems with multiple resource acquisitions and flexible routings
Jonghun Park and Spyros A Reveliotis · 2001
Earlier work this paper cites.
Deep Blue
Murray Campbell, A Joseph Hoane Jr, and Feng-Hsiung Hsu · 2002
Earlier work this paper cites.
On the global convergence of stochastic fictitious play
Josef Hofbauer and William H Sandholm · 2002
Earlier work this paper cites.
Diplomacy artificial intelligence development environment
Andrew Rose, David Normal, and Hamish Williams · 2002
Earlier work this paper cites.
Learning a game strategy using pattern-weights and self-play
Ari Shapiro, Gil Fuchs, and Robert Levinson · 2002
Earlier work this paper cites.
Evolutionary game dynamics
Josef Hofbauer and Karl Sigmund · 2003
Cited alongside, same era.
3-NASH is PPAD-Complete
X. Chen and X. Deng · 2005
Cited alongside, same era.
Three-player games are hard
Constantinos Daskalakis and Christos H. Papadimitriou · 2005
Cited alongside, same era.
General game playing: Overview of the AAAI competition
Michael Genesereth, Nathaniel Love, and Barney Pell · 2005
Cited alongside, same era.
Tactical coordination in no-press Diplomacy
Stefan J Johansson and Fredrik Håård · 2005
Cited alongside, same era.
Generalised weakened fictitious play
David S Leslie and Edmund J Collins · 2006
Cited alongside, same era.
The Colonel Blotto game
Brian Roberson · 2006
Deepstack: Expert-level artificial intelligence in heads-up no-limit Poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Re-evaluating evaluation
David Balduzzi, Karl Tuyls, Julien Perolat, and Thore Graepel · 2018
Later among the works it cites.
Relational inductive biases, deep learning, and graph networks
Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al · 2018
Later among the works it cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trueskill™: a Bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel · 2007
Cited alongside, same era.
Multi-robot perimeter patrol in adversarial settings
Noa Agmon, Sarit Kraus, and Gal A Kaminka · 2008
Cited alongside, same era.
No-regret learning in convex games
Geoffrey J Gordon, Amy Greenwald, and Casey Marks · 2008
Cited alongside, same era.
A testbed for multiagent systems Technical Report IIIA-TR-2009-09
Angela Fabregues and Carles Sierra · 2009
Cited alongside, same era.
Learning and equilibrium
D. Fudenberg and D. Levine · 2009
Cited alongside, same era.
Cooperating with machines
Jacob W Crandall, Mayada Oudah, Fatimah Ishowo-Oloko, Sherief Abdallah, Jean-François Bonnefon, Manuel Cebrian, Azim Shariff, Michael A Goodrich, Iyad Rahwan, et al · 2018
Later among the works it cites.
The challenge of negotiation in the game of Diplomacy
Dave De Jonge, Tim Baarslag, Reyhan Aydoğan, Catholijn Jonker, Katsuhide Fujita, and Takayuki Ito · 2018
Later among the works it cites.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Later among the works it cites.
Learning with opponent-learning awareness
Jakob Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2018
Later among the works it cites.
Bayesian action decoder for deep multi-agent reinforcement learning
Jakob N. Foerster, H. Francis Song, Edward Hughes, Neil Burch, Iain Dunning, Shimon Whiteson, Matthew Botvinick, and Michael Bowling · 2018
Later among the works it cites.
Pommerman: A multi-agent playground
Cinjon Resnick, Wes Eldridge, David Ha, Denny Britz, Jakob Foerster, Julian Togelius, Kyunghyun Cho, and Joan Bruna · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters Chess, Shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2019
Later among the works it cites.
Open-ended learning in symmetric zero-sum games
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Perolat, Max Jaderberg, and Thore Graepel · 2019
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning, 2019
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, Rafal Józefowicz, Scott Gray, Catherine Olsson, Jakub Pachocki, Michael Petrov, Henrique Pondé de Oliveira Pinto, Jonathan Raiman, Tim Salimans, Jeremy Schlatter, Jonas Schneider, Szymon Sidor, Ilya Sutskever, Jie Tang, Filip Wolski, and Susan Zhang · 2019
Later among the works it cites.
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Later among the works it cites.
On the utility of learning about humans for human-AI coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan · 2019
Later among the works it cites.
Learning to correlate in multi-player general-sum sequential games
Andrea Celli, Alberto Marchesi, Tommaso Bianchi, and Nicola Gatti · 2019
Later among the works it cites.
The imitation game: Learned reciprocity in markov games
Tom Eccles, Edward Hughes, János Kramár, Steven Wheelwright, and Joel Z Leibo · 2019
Later among the works it cites.
Coarse correlation in extensive-form games
Gabriele Farina, Tommaso Bianchi, and Tuomas Sandholm · 2019
Later among the works it cites.
The MineRL competition on sample efficient reinforcement learning using human priors
William H Guss, Cayden Codel, Katja Hofmann, Brandon Houghton, Noboru Kuno, Stephanie Milani, Sharada Mohanty, Diego Perez Liebana, Ruslan Salakhutdinov, Nicholay Topin, et al · 2019
Later among the works it cites.
Simplified action decoder for deep multi-agent reinforcement learning, 2019
Hengyuan Hu and Jakob N Foerster · 2019
Later among the works it cites.
Human-level performance in 3D multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castaneda, Charles Beattie, Neil C Rabinowitz, Ari S Morcos, Avraham Ruderman, et al · 2019
Later among the works it cites.
OpenSpiel: A framework for reinforcement learning in games
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Pérolat, Sriram Srinivasan, Finbarr Timbers, Karl Tuyls, Shayegan Omidshafiei, Daniel Hennes, Dustin Morrill, Paul Muller, Timo Ewalds, Ryan Faulkner, János Kramár, Bart De Vylder, Brennan Saeta, James Bradbury, David Ding, Sebastian Borgeaud, Matthew Lai, Julian Schrittwieser, Thomas Anthony, Edward Hughes, Ivo Danihelka, and Jonah Ryan-Davis · 2019
Later among the works it cites.
Improving policies via search in cooperative partially observable games
Adam Lerer, Hengyuan Hu, Jakob Foerster, and Noam Brown · 2019
Later among the works it cites.
Emergent coordination through competition
Siqi Liu, Guy Lever, Josh Merel, Saran Tunyasuvunakool, Nicolas Heess, and Thore Graepel · 2019
Later among the works it cites.
Computing approximate equilibria in sequential adversarial games by exploitability descent
Edward Lockhart, Marc Lanctot, Julien Pérolat, Jean-Baptiste Lespiau, Dustin Morrill, Finbarr Timbers, and Karl Tuyls · 2019
Later among the works it cites.
No-press Diplomacy: Modeling multi-agent gameplay
Philip Paquette, Yuchen Lu, Seton Steven Bocco, Max Smith, O-G Satya, Jonathan K Kummerfeld, Joelle Pineau, Satinder Singh, and Aaron C Courville · 2019
Later among the works it cites.
Finding friend and foe in multi-agent games
Jack Serrino, Max Kleiman-Weiner, David C Parkes, and Josh Tenenbaum · 2019
Later among the works it cites.
Arena: A general evaluation platform and building toolkit for multi-agent intelligence
Yuhang Song, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu, Mai Xu, Zihan Ding, and Lianlong Wu · 2019
Later among the works it cites.
Alphastar: Mastering the real-time strategy game Starcraft II
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Max Jaderberg, Wojciech M Czarnecki, Andrew Dudzik, Aja Huang, Petko Georgiev, Richard Powell, et al · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
POET: open-ended coevolution of environments and their optimized solutions
Rui Wang, Joel Lehman, Jeff Clune, and Kenneth O Stanley · 2019
Later among the works it cites.
Smooth markets: A basic mechanism for organizing gradient-based learners
David Balduzzi, Wojiech M Czarnecki, Thomas W Anthony, Ian M Gemp, Edward Hughes, Joel Z Leibo, Georgios Piliouras, and Thore Graepel · 2020
Closest in time.
The Hanabi challenge: A new frontier for AI research
Nolan Bard, Jakob N Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, et al · 2020
Closest in time.
Learning to resolve alliance dilemmas in many-player zero-sum games
Edward Hughes, Thomas W Anthony, Tom Eccles, Joel Z Leibo, David Balduzzi, and Yoram Bachrach · 2020
Closest in time.
Diplomacy adjudicator test cases
Lucas B. Kruijswijk · 2020
Closest in time.
The math of adjudication
Lucas B. Kruijswijk · 2020
Closest in time.
Diplomacy adjudicator test cases - webdiplomacy
Kestas Kuliukas · 2020
Closest in time.
From Poincaré recurrence to convergence in imperfect information games: Finding equilibrium via regularization, 2020
Julien Perolat, Remi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland, Pedro Ortega, Neil Burch, Thomas Anthony, David Balduzzi, Bart De Vylder, Georgios Piliouras, Marc Lanctot, and Karl Tuyls · 2020
Closest in time.