Fetching the paper…
Reading the bibliography…
No-press Diplomacy is a complex strategy game involving both cooperation and competition that has served as a benchmark for multi-agent AI research.
Zur theorie der gesellschaftsspiele
J v Neumann · 1928
Earlier work this paper cites.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
An analog of the minimax theorem for vector payoffs
David Blackwell et al · 1956
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
The weighted majority algorithm
Nick Littlestone and Manfred K Warmuth · 1994
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Sergiu Hart and Andreu Mas-Colell · 2000
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
Brendan McMahan, Geoffrey Gordon, and Avrim Blum · 2003
Earlier work this paper cites.
Mm algorithms for generalized bradley-terry models
David R. Hunter · 2004
Earlier work this paper cites.
Bayeselo
Rémi Coulom · 2005
Earlier work this paper cites.
Heads-up limit hold’em poker is solved
Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin · 2015
Cited alongside, same era.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2017
Cited alongside, same era.
Dynamic thresholding and pruning for regret minimization
Noam Brown, Christian Kroer, and Tuomas Sandholm · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
To whom tribute is due: The next step in scoring systems, 2020
Brandon Fogel · 2020
Later among the works it cites.
Human-level performance in no-press diplomacy via equilibrium search
Jonathan Gray, Adam Lerer, Anton Bakhtin, and Noam Brown · 2020
Later among the works it cites.
“other-play” for zero-shot coordination
Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob Foerster · 2020
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Noah Y Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Cited alongside, same era.
Learning existing social conventions via observationally augmented self-play
Adam Lerer and Alexander Peysakhovich · 2019
Cited alongside, same era.
No-press diplomacy: Modeling multi-agent gameplay
Philip Paquette, Yuchen Lu, Seton Steven Bocco, Max Smith, O-G Satya, Jonathan K Kummerfeld, Joelle Pineau, Satinder Singh, and Aaron C Courville · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Learning to play no-press diplomacy with best response policy iteration
Thomas Anthony, Tom Eccles, Andrea Tacchetti, János Kramár, Ian Gemp, Thomas Hudson, Nicolas Porcel, Marc Lanctot, Julien Perolat, Richard Everett, Satinder Singh, Thore Graepel, and Yoram Bachrach · 2020
Cited alongside, same era.
No-press diplomacy from scratch
Anton Bakhtin, David Wu, Adam Lerer, and Noam Brown · 2021
Later among the works it cites.
K-level reasoning for zero-shot coordination in hanabi
Brandon Cui, Hengyuan Hu, Luis Pineda, and Jakob Foerster · 2021
Later among the works it cites.
Off-belief learning
Hengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda, David Wu, Noam Brown, and Jakob Foerster · 2021
Later among the works it cites.
Evaluation of human-ai teams for learned and rule-based agents in hanabi
Ho Chit Siu, Jaime Peña, Edenna Chen, Yutai Zhou, Victor Lopez, Kyle Palko, Kimberlee Chang, and Ross Allen · 2021
Later among the works it cites.
Collaborating with humans without human data
DJ Strouse, Kevin McKee, Matt Botvinick, Edward Hughes, and Richard Everett · 2021
Later among the works it cites.
Modeling strong and human-like gameplay with kl-regularized search
Athul Paul Jacob, David J Wu, Gabriele Farina, Adam Lerer, Hengyuan Hu, Anton Bakhtin, Jacob Andreas, and Noam Brown · 2022
Closest in time.