Fetching the paper…
Reading the bibliography…
Securing coordination between AI agent and teammates (human players or AI agents) in contexts involving unfamiliar humans continues to pose a significant challenge in Zero-Shot Coordination.
Cores of convex games
Shapley, L. S. (1971) · 1971
Earlier work this paper cites.
Centrality in social networks conceptual clarification
Freeman, L. C. (1978) · 1978
Earlier work this paper cites.
Td-gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G. (1994) · 1994
Earlier work this paper cites.
The Theory of Learning in Games
Fudenberg, D., & Levine, D. K. (1998) · 1998
Earlier work this paper cites.
The pagerank citation ranking: Bringing order to the web.
Page, L., Brin, S., Motwani, R., & Winograd, T. (1999) · 1999
Earlier work this paper cites.
Coordinated reinforcement learning
Guestrin, C., Lagoudakis, M., & Parr, R. (2002) · 2002
Earlier work this paper cites.
Analyzing complex strategic interactions in multi-agent systems
Walsh, W. E., Das, R., Tesauro, G., & Kephart, J. O. (2002) · 2002
Earlier work this paper cites.
Weighted pagerank algorithm
Xing, W., & Ghorbani, A. (2004) · 2004
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
Legg, S., & Hutter, M. (2007) · 2007
Earlier work this paper cites.
Introduction to the theory of cooperative games
Peleg, B., & Sudhölter, P. (2007) · 2007
Earlier work this paper cites.
Polynomial calculation of the shapley value based on sampling
Castro, J., Gómez, D., & Tejada, J. (2009) · 2009
Earlier work this paper cites.
Computational aspects of cooperative game theory
Chalkiadakis, G., Elkind, E., & Wooldridge, M. (2011) · 2011
Earlier work this paper cites.
Continually adding self-invented problems to the repertoire: First experiments with powerplay.
Srivastava, R., Steunebrink, B., Stollenga, M., & Schmidhuber, J. (2012) · 2012
Earlier work this paper cites.
Population based training of neural networks.
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., Fernando, C., & Kavukcuoglu, K. (2017) · 2017
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Perolat, J., Silver, D., & Graepel, T. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms.
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017) · 2017
Earlier work this paper cites.
Learning social conventions in Markov games.
Lerer, A., & Peysakhovich, A. (2018) · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., & Hassabis, D. (2018) · 2018
Cited alongside, same era.
A generalised method for empirical game theoretic analysis
Tuyls, K., Pérolat, J., Lanctot, M., Leibo, J. Z., & Graepel, T. (2018) · 2018
Cited alongside, same era.
Open-ended learning in symmetric zero-sum games
Balduzzi, D., Garnelo, M., Bachrach, Y., Czarnecki, W., Pérolat, J., Jaderberg, M., & Graepel, T. (2019) · 2019
Cited alongside, same era.
Evaluating fluency in human–robot collaboration
Hoffman, G. (2019) · 2019
Cited alongside, same era.
The hanabi challenge: A new frontier for ai research
Bard, N., Foerster, J. N., Chandar, S., Burch, N., Lanctot, M., Song, H. F., Parisotto, E., Dumoulin, V., Moitra, S., Hughes, E., et al. (2020) · 2020
Cited alongside, same era.
Collaborating with humans without human data
Strouse, D., McKee, K., Botvinick, M., Hughes, E., & Everett, R. (2021) · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
Team, O.-E. L., Stooke, A., Mahajan, A., Barros, C., Deck, C., Bauer, J., Sygnowski, J., Trebacz, M., Jaderberg, M., Mathieu, M., McAleese, N., Bradley-Schmieg, N., Wong, N., Porcel, N., Raileanu, R., Hughes-Fitt, S., Dalibard, V., & Czarnecki, W. M. (2021) · 2021
Later among the works it cites.
Diverse auto-curriculum is critical for successful real-world multiagent learning systems.
Yang, Y., Luo, J., Wen, Y., Slumbers, O., Graves, D., Bou Ammar, H., Wang, J., & Taylor, M. E. (2021) · 2021
Later among the works it cites.
Maximum entropy population based training for zero-shot Human-AI coordination.
Zhao, R., Song, J., Haifeng, H., Gao, Y., Wu, Y., Sun, Z., & Wei, Y. (2021) · 2021
Later among the works it cites.
Generating and adapting to diverse ad-hoc partners in hanabi.
Canaan, R., Gao, X., Togelius, J., Nealen, A., & Menzel, S. (2022) · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep coordination graphs
Böhmer, W., Kurin, V., & Whiteson, S. (2020) · 2020
Cited alongside, same era.
On the utility of learning about humans for human-ai coordination
Carroll, M., Shah, R., Ho, M. K., Griffiths, T. L., Seshia, S. A., Abbeel, P., & Dragan, A. (2020) · 2020
Cited alongside, same era.
Investigating partner diversification methods in cooperative multi-agent deep reinforcement learning
Charakorn, R., Manoonpong, P., & Dilokthanakul, N. (2020) · 2020
Cited alongside, same era.
At your service: Coffee beans recommendation from a robot assistant.
de Berardinis, J., Pizzuto, G., Lanza, F., Chella, A., Meira, J., & Cangelosi, A. (2020) · 2020
Cited alongside, same era.
"other-play" for zero-shot coordination
Hu, H., Lerer, A., Peysakhovich, A., & Foerster, J. (2020) · 2020
Cited alongside, same era.
Pipeline psro: A scalable approach for finding approximate nash equilibria in large games
McAleer, S., Lanier, J., Fox, R., & Baldi, P. (2020) · 2020
Cited alongside, same era.
Towards playing full MOBA games with deep reinforcement learning
Ye, D., Chen, G., Zhang, W., Chen, S., Yuan, B., Liu, B., Chen, J., Liu, Z., Qiu, F., Yu, H., Yin, Y., Shi, B., Wang, L., Shi, T., Fu, Q., Yang, W., Huang, L., & Liu, W. (2020) · 2020
Cited alongside, same era.
Later among the works it cites.
Zero-shot assistance in novel decision problems.
De Peuter, S., & Kaski, S. (2022) · 2022
Later among the works it cites.
Generalization in cooperative multi-agent systems.
Mahajan, A., Samvelyan, M., Gupta, T., Ellis, B., Sun, M., Rocktäschel, T., & Whiteson, S. (2022) · 2022
Later among the works it cites.
Self-play psro: Toward optimal populations in two-player zero-sum games.
McAleer, S., Lanier, J., Wang, K., Baldi, P., Fox, R., & Sandholm, T. (2022) · 2022
Later among the works it cites.
Open-ended reinforcement learning with neural reward functions
Meier, R., & Mujika, A. (2022) · 2022
Later among the works it cites.
A general learning framework for open ad hoc teamwork using graph-based policy learning.
Rahman, A., Carlucho, I., Höpner, N., & Albrecht, S. V. (2022) · 2022
Later among the works it cites.
Honor of kings arena: an environment for generalization in competitive reinforcement learning
Wei, H., Chen, J., Ji, X., Qin, H., Deng, M., Li, S., Wang, L., Zhang, W., Yu, Y., Lin, L., Huang, L., Ye, D., Fu, Q., & Yang, W. (2022) · 2022
Later among the works it cites.
Heterogeneous multi-agent zero-shot coordination by coevolution.
Xue, K., Wang, Y., Yuan, L., Guan, C., Qian, C., & Yu, Y. (2022) · 2022
Later among the works it cites.
Generating diverse cooperative agents by learning incompatible policies
Charakorn, R., Manoonpong, P., & Dilokthanakul, N. (2023) · 2023
Closest in time.
Towards effective and interpretable human-agent collaboration in MOBA games: A communication perspective
Gao, Y., Liu, F., Wang, L., Lian, Z., Wang, W., Li, S., Wang, X., Zeng, X., Wang, R., Wang, J., Fu, Q., Yang, W., Huang, L., & Liu, W. (2023) · 2023
Closest in time.
Segment anything.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023) · 2023
Closest in time.
PECAN: leveraging policy ensemble for context-aware zero-shot human-ai coordination
Lou, X., Guo, J., Zhang, J., Wang, J., Huang, K., & Du, Y. (2023) · 2023
Closest in time.
Learning zero-shot cooperation with humans, assuming humans are biased
Yu, C., Gao, J., Liu, W., Xu, B., Tang, H., Yang, J., Wang, Y., & Wu, Y. (2023) · 2023
Closest in time.