Fetching the paper…
Reading the bibliography…
Training agents in cooperative settings offers the promise of AI agents able to interact effectively with humans (and other agents) in the real world.
Does the chimpanzee have a theory of mind?
D. Premack and G. Woodruff · 1978
Earlier work this paper cites.
Stochastic relaxation, Gibbs distributions and Bayesian restoration of images
S. Geman and D. Geman · 1984
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
G. Tesauro · 1994
Earlier work this paper cites.
Evolutionary Computation 1: Basic Algorithms and Operators
T. Baeck, D. Fogel, and Z. Michalewicz · 2000
Earlier work this paper cites.
Minimum-entropy data partitioning using reversible jump Markov chain Monte Carlo
S. J. Roberts, C. Holmes, and D. Denison · 2001
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein · 2002
Earlier work this paper cites.
Taming decentralized POMDPs: Towards efficient policy computation for multiagent settings
R. Nair, M. Tambe, M. Yokoo, D. Pynadath, and S. Marsella · 2003
Earlier work this paper cites.
Jags: A program for analysis of bayesian graphical models using gibbs sampling
M. Plummer et al · 2003
Earlier work this paper cites.
Coordination and adaptation in impromptu teams
M. Bowling and P. McCracken · 2005
Earlier work this paper cites.
Pattern recognition and machine learning
C. M. Bishop · 2006
Earlier work this paper cites.
From external to internal regret
A. Blum and Y. Mansour · 2007
Earlier work this paper cites.
Awesome: A general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents
V. Conitzer and T. Sandholm · 2007
Earlier work this paper cites.
Openbugs user manual
D. Spiegelhalter, A. Thomas, N. Best, and D. Lunn · 2007
Earlier work this paper cites.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2007
Earlier work this paper cites.
E. Brochu, V. M. Cora, and N. de Freitas · 2010
Earlier work this paper cites.
Pymc: Bayesian stochastic modelling in python
A. Patil, D. Huard, and C. J. Fonnesbeck · 2010
Earlier work this paper cites.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
P. Stone, G. A. Kaminka, S. Kraus, and J. S. Rosenschein · 2010
Earlier work this paper cites.
Learning to compete, coordinate, and cooperate in repeated games using reinforcement learning
J. W. Crandall and M. A. Goodrich · 2011
Earlier work this paper cites.
Online bandit learning against an adaptive adversary: from regret to policy regret
R. Arora, O. Dekel, and A. Tewari · 2012
Earlier work this paper cites.
Online implicit agent modelling
N. Bard, M. Johanson, N. Burch, and M. Bowling · 2013
Earlier work this paper cites.
R. Gibson · 2013
Earlier work this paper cites.
A compilation target for probabilistic programming languages
B. Paige and F. Wood · 2014
Earlier work this paper cites.
Illuminating search spaces by mapping elites
J.-B. Mouret and J. Clune · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson · 2016
Cited alongside, same era.
4. the bargaining problem
J. F. Nash · 2016
Cited alongside, same era.
Learning multiagent communication with backpropagation
S. Sukhbaatar, A. Szlam, and R. Fergus · 2016
Cited alongside, same era.
Rational quantitative attribution of beliefs, desires and percepts in human mentalizing
C. L. Baker, J. Jara-Ettinger, R. Saxe, and J. B. Tenenbaum · 2017
Cited alongside, same era.
Making friends on the fly: Cooperating with new teammates
S. Barrett, A. Rosenfeld, S. Kraus, and P. Stone · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch · 2017
Cited alongside, same era.
Accelerating bayesian inference on structured graphs using parallel gibbs sampling
G. G. Ko, Y. Chai, R. A. Rutenbar, D. Brooks, and G.-Y. Wei · 2019
Later among the works it cites.
Flexgibbs: Reconfigurable parallel gibbs sampling accelerator for structured graphs
G. G. Ko, Y. Chai, R. A. Rutenbar, D. Brooks, and G.-Y. Wei · 2019
Later among the works it cites.
Learning existing social conventions via observationally augmented self-play
A. Lerer and A. Peysakhovich · 2019
Later among the works it cites.
Multi-agent adversarial inverse reinforcement learning
L. Yu, J. Song, and S. Ermon · 2019
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Y. Bai and C. Jin · 2020
Later among the works it cites.
The Hanabi challenge: A new frontier for AI research
N. Bard, J. N. Foerster, S. Chandar, N. Burch, M. Lanctot, H. F. Song, E. Parisotto, V. Dumoulin, S. Moitra, E. Hughes, et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
M. Moravčík, M. Schmid, N. Burch, V. Lisỳ, D. Morrill, N. Bard, T. Davis, K. Waugh, M. Johanson, and M. Bowling · 2017
Cited alongside, same era.
Prosocial learning agents solve generalized stag hunts better than selfish ones
A. Peysakhovich and A. Lerer · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Cited alongside, same era.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
N. Brown and T. Sandholm · 2018
Cited alongside, same era.
Improving exploration in Evolution Strategies for deep reinforcement learning via a population of novelty-seeking agents
E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. O. Stanley, and J. Clune · 2018
Cited alongside, same era.
Later among the works it cites.
Evaluating rl agents in hanabi with unseen partners
R. Canaan, X. Gao, Y. Chung, J. Togelius, A. Nealen, and S. Menzel · 2020
Later among the works it cites.
Open problems in cooperative AI
A. Dafoe, E. Hughes, Y. Bachrach, T. Collins, K. R. McKee, J. Z. Leibo, K. Larson, and T. Graepel · 2020
Later among the works it cites.
Slime volleyball gym environment
D. Ha · 2020
Later among the works it cites.
Simplified action decoder for deep multi-agent reinforcement learning
H. Hu and J. N. Foerster · 2020
Later among the works it cites.
“other-play”for zero-shot coordination
H. Hu, A. Peysakhovich, A. Lerer, and J. Foerster · 2020
Later among the works it cites.
Improving policies via search in cooperative partially observable games
A. Lerer, H. Hu, J. N. Foerster, and N. Brown · 2020
Later among the works it cites.
Database-independent molecular formula annotation using gibbs sampling through zodiac
M. Ludwig, L.-F. Nothias, K. Dührkop, I. Koester, M. Fleischauer, M. A. Hoffmann, D. Petras, F. Vargas, M. Morsy, L. Aluwihare, et al · 2020
Later among the works it cites.
Ridge Rider: Finding diverse solutions by following eigenvectors of the Hessian
J. Parker-Holder, L. Metz, C. Resnick, H. Hu, A. Lerer, A. Letcher, A. Peysakhovich, A. Pacchiano, and J. Foerster · 2020
Later among the works it cites.
Effective diversity in population based reinforcement learning
J. Parker-Holder, A. Pacchiano, K. M. Choromanski, and S. J. Roberts · 2020
Later among the works it cites.
Multi-agent determinantal Q-learning
Y. Yang, Y. Wen, J. Wang, L. Chen, K. Shao, D. Mguni, and W. Zhang · 2020
Later among the works it cites.
Survey of self-play in reinforcement learning
A. DiGiovanni and E. C. Zell · 2021
Later among the works it cites.
Off-belief learning
H. Hu, A. Lerer, B. Cui, L. Pineda, D. Wu, N. Brown, and J. N. Foerster · 2021
Later among the works it cites.
Trajectory diversity for zero-shot coordination
A. Lupu, H. Hu, and J. Foerster · 2021
Later among the works it cites.
Continuous coordination as a realistic scenario for lifelong learning, 2021
H. Nekoei, A. Badrinaaraayanan, A. Courville, and S. Chandar · 2021
Later among the works it cites.
HOAD: The Hanabi open agent dataset
A. Sarmasi, T. Zhang, C.-H. Cheng, H. Pham, X. Zhou, D. Nguyen, S. Shekdar, and J. McCoy · 2021
Later among the works it cites.
Collaborating with humans without human data
D. Strouse, K. R. McKee, M. Botvinick, E. Hughes, and R. Everett · 2021
Later among the works it cites.
The surprising effectiveness of MAPPO in cooperative, multi-agent games
C. Yu, A. Velu, E. Vinitsky, Y. Wang, A. M. Bayen, and Y. Wu · 2021
Later among the works it cites.