Fetching the paper…
Reading the bibliography…
Zero-shot coordination (ZSC) is a new cooperative multi-agent reinforcement learning (MARL) challenge that aims to train an ego agent to work with diverse, unseen partners during deployment.
The proof and measurement of association between two things
Charles Spearman · 1961
Earlier work this paper cites.
Stochastic dynamic programming: successive approximations and nearly optimal strategies for markov decision processes and markov games
Johannes Van Der Wal · 1980
Earlier work this paper cites.
Principal components analysis , volume 69
George H Dunteman · 1989
Earlier work this paper cites.
A course in game theory
Martin J Osborne and Ariel Rubinstein · 1994
Earlier work this paper cites.
Td-gammon, a self-teaching backgammon program, achieves master-level play
Gerald Tesauro · 1994
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart Russell, et al · 2000
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
Daniel S. Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein · 2002
Earlier work this paper cites.
The reward hypothesis, 2004
Richard S Sutton and Andrew G Barto · 2004
Earlier work this paper cites.
Ad hoc autonomous agent teams: collaboration without pre-coordination
Peter Stone, Gal A. Kaminka, Sarit Kraus, and Jeffrey S. Rosenschein · 2010
Earlier work this paper cites.
Determinantal point processes for machine learning
Alex Kulesza, Ben Taskar, et al · 2012
Earlier work this paper cites.
Three years of the robocup standard platform league drop-in player competition: Creating and maintaining a large scale ad hoc teamwork robotics competition
Katie Genter, Tim Laue, and Peter Stone · 2017
Earlier work this paper cites.
Population based training of neural networks, 2017
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M. Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, Chrisantha Fernando, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Variational inverse control with events: A general framework for data-driven reward definition
Justin Fu, Avi Singh, Dibya Ghosh, Larry Yang, and Sergey Levine · 2018
Earlier work this paper cites.
Deep reinforcement learning for event-driven multi-agent decision processes
Kunal Menda, Yi-Chun Chen, Justin Grana, James W. Bono, Brendan D. Tracey, Mykel J. Kochenderfer, and David Wolpert · 2018
Earlier work this paper cites.
Time limits in reinforcement learning
Fabio Pardo, Arash Tavakoli, Vitaly Levdik, and Petar Kormushev · 2018
Earlier work this paper cites.
Socialrobot: Towards a personalized elderly care mobile robot
David Portugal, Luís Santos, Pedro Trindade, Christophoros Christophorou, Panayiotis Andreou, Dimosthenis Georgiadis, Marios Belk, João Freire, Paulo Alvito, George Samaras, Eleni Christodoulou, and Jorge Dias · 2018
Earlier work this paper cites.
On the utility of learning about humans for human-ai coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan · 2019
Earlier work this paper cites.
"other-play" for zero-shot coordination
Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob N. Foerster · 2020
Earlier work this paper cites.
Pytorch implementations of reinforcement learning algorithms
Ilya Kostrikov and Matin Raayai Ardakani · 2020
Cited alongside, same era.
Google research football: A novel reinforcement learning environment
Karol Kurach, Anton Raichuk, Piotr Stańczyk, Michał Zając, Olivier Bachem, Lasse Espeholt, Carlos Riquelme, Damien Vincent, Marcin Michalski, Olivier Bousquet, et al · 2020
Cited alongside, same era.
Effective diversity in population based reinforcement learning
Jack Parker-Holder, Aldo Pacchiano, Krzysztof M Choromanski, and Stephen J Roberts · 2020
Cited alongside, same era.
Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder De Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville, and Marc G. Bellemare · 2021
Cited alongside, same era.
The surprising effectiveness of ppo in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu · 2022
Later among the works it cites.
Picor: Multi-task deep reinforcement learning with policy correction
Fengshuo Bai, Hongming Zhang, Tianyang Tao, Zhiheng Wu, Yanna Wang, and Bo Xu · 2023
Closest in time.
Settling the reward hypothesis
Michael Bowling, John D. Martin, David Abel, and Will Dabney · 2023
Closest in time.
Generating diverse cooperative agents by learning incompatible policies
Rujikorn Charakorn, Poramate Manoonpong, and Nat Dilokthanakul · 2023
Closest in time.
A review of cooperation in multi-agent learning
Yali Du, Joel Z Leibo, Usman Islam, Richard Willis, and Peter Sunehag · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Evaluating the robustness of collaborative agents
Paul Knott, Micah Carroll, Sam Devlin, Kamil Ciosek, Katja Hofmann, Anca D. Dragan, and Rohin Shah · 2021
Cited alongside, same era.
Towards out-of-distribution generalization: A survey
Jiashuo Liu, Zheyan Shen, Yue He, Xingxuan Zhang, Renzhe Xu, Han Yu, and Peng Cui · 2021
Cited alongside, same era.
Trajectory diversity for zero-shot coordination
Andrei Lupu, Brandon Cui, Hengyuan Hu, and Jakob N. Foerster · 2021
Cited alongside, same era.
Bootstrap confidence intervals, 2021
E. Helwig Nathaniel · 2021
Cited alongside, same era.
Evaluation of human-ai teams for learned and rule-based agents in hanabi
Ho Chit Siu, Jaime Daniel Peña, Edenna Chen, Yutai Zhou, Victor J. Lopez, Kyle Palko, Kimberlee C. Chang, and Ross E. Allen · 2021
Cited alongside, same era.
Collaborating with humans without human data
DJ Strouse, Kevin McKee, Matt Botvinick, Edward Hughes, and Richard Everett · 2021
Cited alongside, same era.
Model-based multi-agent policy optimization with adaptive opponent-wise rollouts
Weinan Zhang, Xihuai Wang, Jian Shen, and Ming Zhou · 2021
Cited alongside, same era.
Dom Huh and Prasant Mohapatra · 2023
Closest in time.
A survey of zero-shot generalisation in deep reinforcement learning
Robert Kirk, Amy Zhang, Edward Grefenstette, and Tim Rocktäschel · 2023
Closest in time.
Who needs to know? Minimal knowledge for optimal coordination
Niklas Lauffer, Ameesh Shah, Micah Carroll, Michael D Dennis, and Stuart Russell · 2023
Closest in time.
A hierarchical approach to population training for human-ai collaboration
Yi Loo, Chen Gong, and Malika Meghjani · 2023
Closest in time.
Pecan: Leveraging policy ensemble for context-aware zero-shot human-ai coordination
Xingzhou Lou, Jiaxian Guo, Junge Zhang, Jun Wang, Kaiqi Huang, and Yali Du · 2023
Closest in time.
Diverse conventions for human-AI collaboration
Bidipta Sarkar, Andy Shih, and Dorsa Sadigh · 2023
Closest in time.
An efficient end-to-end training approach for zero-shot human-ai coordination
Xue Yan, Jiaxian Guo, Xingzhou Lou, Jun Wang, Haifeng Zhang, and Yali Du · 2023
Closest in time.
Learning zero-shot cooperation with humans, assuming humans are biased
Chao Yu, Jiaxuan Gao, Weilin Liu, Botian Xu, Hao Tang, Jiaqi Yang, Yu Wang, and Yi Wu · 2023
Closest in time.
A survey of progress on cooperative multi-agent reinforcement learning in open environment
Lei Yuan, Ziqian Zhang, Lihe Li, Cong Guan, and Yang Yu · 2023
Closest in time.
Deep long-tailed learning: A survey
Yifan Zhang, Bingyi Kang, Bryan Hooi, Shuicheng Yan, and Jiashi Feng · 2023
Closest in time.
Maximum entropy population-based training for zero-shot human-ai coordination
Rui Zhao, Jinming Song, Yufeng Yuan, Haifeng Hu, Yang Gao, Yi Wu, Zhongqian Sun, and Wei Yang · 2023
Closest in time.
Tayuki Osa and Tatsuya Harada · 2024
Closest in time.
Yiwei Shi, Muning Wen, Qi Zhang, Weinan Zhang, Cunjia Liu, and Weiru Liu · 2024
Closest in time.
Mutual theory of mind in human-ai collaboration: An empirical study with llm-driven ai agents in a real-time shared workspace task
Shao Zhang*, Xihuai Wang*, Wenhao Zhang, Yongshan Chen, Landi Gao, Dakuo Wang, Weinan Zhang, Xinbing Wang, and Ying Wen · 2024
Closest in time.