Fetching the paper…
Reading the bibliography…
Cooperative Multi-agent Reinforcement Learning (MARL) algorithms with Zero-Shot Coordination (ZSC) have gained significant attention in recent years.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
Peter Stone, Gal A Kaminka, Sarit Kraus, and Jeffrey S Rosenschein · 2000
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
Daniel S Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein · 2002
Earlier work this paper cites.
" other-play" for zero-shot coordination
Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob Foerster · 2003
Earlier work this paper cites.
Taming decentralized pomdps: Towards efficient policy computation for multiagent settings
Ranjit Nair, Milind Tambe, Makoto Yokoo, David Pynadath, and Stacy Marsella · 2003
Earlier work this paper cites.
Coordination and adaptation in impromptu teams
Michael Bowling and Peter McCracken · 2005
Earlier work this paper cites.
Eric Brochu, Vlad M Cora, and Nando de Freitas · 2010
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, …, and Stig Petersen · 2015
Earlier work this paper cites.
Illuminating search spaces by mapping elites
Jean-Baptiste Mouret and Jeff Clune · 2015
Earlier work this paper cites.
A concise introduction to decentralized POMDPs
Frans A Oliehoek and Christopher Amato · 2016
Earlier work this paper cites.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Earlier work this paper cites.
Learning multiagent communication with backpropagation
Sainbayar Sukhbaatar, Rob Fergus, et al · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas · 2016
Cited alongside, same era.
Making friends on the fly: Cooperating with new teammates
Samuel Barrett, Avi Rosenfeld, Sarit Kraus, and Peter Stone · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Prosocial learning agents solve generalized stag hunts better than selfish ones
Alexander Peysakhovich and Adam Lerer · 2017
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2017
Evaluation of human-ai teams for learned and rule-based agents in hanabi
Jinkyoo Kim, Hua Yang, Sungwook Lee, and Kyunghyun Cho · 2020
Later among the works it cites.
Google research football: A novel reinforcement learning environment
Karol Kurach, Anton Raichuk, Piotr Stańczyk, Michał Zając, Olivier Bachem, Lasse Espeholt, Carlos Riquelme, Damien Vincent, Marcin Michalski, Olivier Bousquet, et al · 2020
Later among the works it cites.
The loca regret: a consistent metric to evaluate model-based behavior in reinforcement learning
Harm Van Seijen, Hadi Nekoei, Evan Racah, and Sarath Chandar · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Later among the works it cites.
K-level reasoning for zero-shot coordination in hanabi
Brandon Cui, Hengyuan Hu, Luis Pineda, and Jakob Foerster · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distributed prioritized experience replay
Daniel Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, and Hado Van Hasselt · 2018
Cited alongside, same era.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2018
Cited alongside, same era.
Diverse agents for ad-hoc cooperation in hanabi
Remi Canaan, Julian Togelius, Andrew Nealen, and Stephan Menzel · 2019
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
Szymon Kapturowski, Tom Schaul, Bilal Piot, Matteo Hessel, Hado van Hasselt, and Marc Lanctot · 2019
Cited alongside, same era.
Learning existing social conventions via observationally augmented self-play
Adam Lerer and Alexander Peysakhovich · 2019
Cited alongside, same era.
The starcraft multi-agent challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder De Witt, Gregory Farquhar, Nantas Nardelli, Tim GJ Rudner, Chia-Man Hung, Philip HS Torr, Jakob Foerster, and Shimon Whiteson · 2019
Cited alongside, same era.
The hanabi challenge: A new frontier for ai research
Nolan Bard, Jakob N Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, et al · 2020
Cited alongside, same era.
Hengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda, Noam Brown, and Jakob Foerster · 2021
Later among the works it cites.
Trajectory diversity for zero-shot coordination
Andrei Lupu, Brandon Cui, Hengyuan Hu, and Jakob Foerster · 2021
Later among the works it cites.
Continuous coordination as a realistic scenario for lifelong learning
Hadi Nekoei, Akilesh Badrinaaraayanan, Aaron Courville, and Sarath Chandar · 2021
Later among the works it cites.
A new formalism, method and open issues for zero-shot coordination
Johannes Treutlein, Michael Dennis, Caspar Oesterheld, and Jakob Foerster · 2021
Later among the works it cites.
Too many cooks: Bayesian inference for coordinating multi-agent collaboration
Sarah A Wu, Rose E Wang, James A Evans, Joshua B Tenenbaum, David C Parkes, and Max Kleiman-Weiner · 2021
Later among the works it cites.
Any-play: An intrinsic augmentation for zero-shot coordination, 2022
Keane Lucas and Ross E. Allen · 2022
Later among the works it cites.
On-the-fly strategy adaptation for ad-hoc agent coordination
Jaleh Zand, Jack Parker-Holder, and Stephen J Roberts · 2022
Later among the works it cites.