Fetching the paper…
Reading the bibliography…
Cooperation in multi-agent learning (MAL) is a topic at the intersection of numerous disciplines, including game theory, economics, social sciences, and evolutionary biology.
Joel Z Leibo, Edward Hughes, Marc Lanctot, and Thore Graepel · 1903
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
The Strategy of Conflict: with a new Preface by the Author
Thomas C Schelling · 1960
Earlier work this paper cites.
A Theory of Justice
John Rawls · 1971
Earlier work this paper cites.
Hockey helmets, concealed weapons, and daylight saving: A study of binary choices with externalities
Thomas C Schelling · 1973
Earlier work this paper cites.
Prisoner’s dilemma—recollections and observations
Anatol Rapoport · 1974
Earlier work this paper cites.
Effective Choice in the Prisoner’s Dilemma
Robert Axelrod · 1980
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
Craig Boutilier · 1996
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Exact and approximate algorithms for partially observable Markov decision processes
Anthony Rocco Cassandra · 1998
Earlier work this paper cites.
Dynamic noncooperative game theory
Tamer Başar and Geert Jan Olsder · 1998
Earlier work this paper cites.
A unified analysis of value-function-based reinforcement-learning algorithms
Csaba Szepesvári and Michael L Littman · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis · 1999
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
Satinder Singh, Tommi Jaakkola, Michael L Littman, and Csaba Szepesvári · 2000
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
Martin Lauer and Martin A Riedmiller · 2000
Earlier work this paper cites.
Friend-or-foe q-learning in general-sum games
Michael L Littman et al · 2001
Earlier work this paper cites.
Optimal payoff functions for members of collectives
David H Wolpert and Kagan Tumer · 2001
Earlier work this paper cites.
Reinforcement learning to play an optimal nash equilibrium in team markov games
Xiaofeng Wang and Tuomas Sandholm · 2002
Earlier work this paper cites.
Qatten: A general framework for cooperative multiagent reinforcement learning
Yaodong Yang, Jianye Hao, Ben Liao, Kun Shao, Guangyong Chen, Wulong Liu, and Hongyao Tang · 2002
Earlier work this paper cites.
Learning dynamics in social dilemmas
Michael W Macy and Andreas Flache · 2002
Earlier work this paper cites.
Analyzing Complex Strategic Interactions in Multi-Agent Systems
William E Walsh, Rajarshi Das, Gerald Tesauro, and Jeffrey O Kephart · 2002
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Earlier work this paper cites.
Methods for Empirical Game-Theoretic Analysis
Michael P Wellman · 2006
Earlier work this paper cites.
Five rules for the evolution of cooperation
Martin A. Nowak · 2006
Earlier work this paper cites.
If multi-agent learning is the answer, what is the question?
Yoav Shoham, Rob Powers, and Trond Grenager · 2007
Earlier work this paper cites.
The network structure of exploration and exploitation
David Lazer and Allan Friedman · 2007
Earlier work this paper cites.
Altruism, selfishness, and spite in traffic routing
Po-An Chen and David Kempe · 2008
Earlier work this paper cites.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
Peter Stone, Gal Kaminka, Sarit Kraus, and Jeffrey Rosenschein · 2010
Earlier work this paper cites.
Lab experiments for the study of social-ecological systems
Marco A Janssen, Robert Holahan, Allen Lee, and Elinor Ostrom · 2010
Earlier work this paper cites.
The origins of the Gini index: Extracts from Variabilità e Mutabilità (1912) by Corrado Gini
Lidia Ceriani and Paolo Verme · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Selfishness level of strategic games
Krzysztof R Apt and Guido Schäfer · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, and Georg Ostrovski · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, and Marc Lanctot · 2016
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando de Freitas, and Shimon Whiteson · 2016
Earlier work this paper cites.
Learning multiagent communication with backpropagation
Sainbayar Sukhbaatar and Rob Fergus · 2016
Earlier work this paper cites.
A concise introduction to decentralized POMDPs , volume 1
Frans A Oliehoek, Christopher Amato, et al · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas · 2016
Earlier work this paper cites.
Social learning strategies modify the effect of network structure on group performance
Daniel Barkoczi and Mirta Galesic · 2016
Earlier work this paper cites.
Coordinate to cooperate or compete: abstract goals and joint intentions in social interaction
Max Kleiman-Weiner, Mark K Ho, Joseph L Austerweil, Michael L Littman, and Joshua B Tenenbaum · 2016
Earlier work this paper cites.
Multi-agent reinforcement learning in sequential social dilemmas
Joel Z Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Stabilising experience replay for deep multi-agent reinforcement learning
Jakob Foerster, Nantas Nardelli, Gregory Farquhar, Triantafyllos Afouras, Philip HS Torr, Pushmeet Kohli, and Shimon Whiteson · 2017
Cited alongside, same era.
A multi-agent reinforcement learning model of common-pool resource appropriation
Julien Perolat, Joel Z Leibo, Vinicius Zambaldi, Charles Beattie, Karl Tuyls, and Thore Graepel · 2017
Cited alongside, same era.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games
Learning to Resolve Alliance Dilemmas in Many-Player Zero-Sum Games
Edward Hughes, Thomas W Anthony, Tom Eccles, Joel Z Leibo, David Balduzzi, and Yoram Bachrach · 2020
Later among the works it cites.
Gifting in multi-agent reinforcement learning
Andrei Lupu and Doina Precup · 2020
Later among the works it cites.
Methodological Individualism
Joseph Heath · 2020
Later among the works it cites.
Inducing cooperation through reward reshaping based on peer evaluations in deep multi-agent reinforcement learning
David Earl Hostallero, Daewoo Kim, Sangwoo Moon, Kyunghwan Son, Wan Ju Kang, and Yung Yi · 2020
Later among the works it cites.
Foolproof Cooperative Learning
Alexis Jacq, Julien Perolat, Matthieu Geist, and Olivier Pietquin · 2020
Later among the works it cites.
Emergent Reciprocity and Team Formation from Randomized Uncertain Social Preferences
Bowen Baker · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peng Peng, Quan Yuan, Ying Wen, Yaodong Yang, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Cited alongside, same era.
Vain: Attentional multi-agent predictive modeling
Yedid Hoshen · 2017
Cited alongside, same era.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Cited alongside, same era.
Starcraft ii: A new challenge for reinforcement learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, and Julian Schrittwieser · 2017
Cited alongside, same era.
Maintaining cooperation in complex social dilemmas using deep reinforcement learning
Adam Lerer and Alexander Peysakhovich · 2017
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Fully decentralized multi-agent reinforcement learning with networked agents
Kaiqing Zhang, Zhuoran Yang, Han Liu, Tong Zhang, and Tamer Basar · 2018
Cited alongside, same era.
Later among the works it cites.
On the utility of learning about humans for human-ai coordination, 2020
Micah Carroll, Rohin Shah, Mark K. Ho, Thomas L. Griffiths, Sanjit A. Seshia, Pieter Abbeel, and Anca Dragan · 2020
Later among the works it cites.
Google research football: A novel reinforcement learning environment
Karol Kurach, Anton Raichuk, Piotr Stańczyk, Michał Zając, Olivier Bachem, Lasse Espeholt, Carlos Riquelme, Damien Vincent, Marcin Michalski, Olivier Bousquet, et al · 2020
Later among the works it cites.
Shared experience actor-critic for multi-agent reinforcement learning
Filippos Christianos, Lukas Schäfer, and Stefano V Albrecht · 2020
Later among the works it cites.
Towards open ad hoc teamwork using graph-based policy learning
Muhammad A Rahman, Niklas Hopner, Filippos Christianos, and Stefano V Albrecht · 2021
Later among the works it cites.
Scalable evaluation of multi-agent reinforcement learning with melting pot
Joel Z Leibo, Edgar A Dueñez-Guzman, Alexander Vezhnevets, John P Agapiou, Peter Sunehag, Raphael Koster, Jayd Matyas, Charlie Beattie, Igor Mordatch, and Thore Graepel · 2021
Later among the works it cites.
Collaborating with humans without human data
DJ Strouse, Kevin McKee, Matt Botvinick, Edward Hughes, and Richard Everett · 2021
Later among the works it cites.
Trust region policy optimisation in multi-agent reinforcement learning
Jakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang · 2021
Later among the works it cites.
Multi-agent graph-attention communication and teaming
Yaru Niu, Rohan Paleja, and Matthew Gombolay · 2021
Later among the works it cites.
Trajectory diversity for zero-shot coordination
Andrei Lupu, Brandon Cui, Hengyuan Hu, and Jakob Foerster · 2021
Later among the works it cites.
Maximum entropy population based training for zero-shot human-ai coordination
Rui Zhao, Jinming Song, Hu Haifeng, Yang Gao, Yi Wu, Zhongqian Sun, and Yang Wei · 2021
Later among the works it cites.
Emergent Prosociality in Multi-Agent Games Through Gifting
Woodrow Z. Wang, Mark Beliaev, Erdem Bıyık, Daniel A. Lazar, Ramtin Pedarsani, and Dorsa Sadigh · 2021
Later among the works it cites.
A multi-agent reinforcement learning model of reputation and cooperation in human groups
Kevin R McKee, Edward Hughes, Tina O Zhu, Martin J Chadwick, Raphael Koster, Antonio Garcia Castaneda, Charlie Beattie, Thore Graepel, Matt Botvinick, and Joel Z Leibo · 2021
Later among the works it cites.
Cooperation and Reputation Dynamics with Reinforcement Learning
Nicolas Anastassacos, Julian García, Stephen Hailes, and Mirco Musolesi · 2021
Later among the works it cites.
Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks
Georgios Papoudakis, Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht · 2021
Later among the works it cites.
Emergent bartering behaviour in multi-agent reinforcement learning
Michael Bradley Johanson, Edward Hughes, Finbarr Timbers, and Joel Z Leibo · 2022
Later among the works it cites.
John P Agapiou, Alexander Sasha Vezhnevets, Edgar A Duéñez-Guzmán, Jayd Matyas, Yiran Mao, Peter Sunehag, Raphael Köster, Udari Madhushani, Kavya Kopparapu, Ramona Comanescu, et al · 2022
Later among the works it cites.
Maser: Multi-agent reinforcement learning with subgoals generated from experience replay buffer
Jeewon Jeon, Woojun Kim, Whiyoung Jung, and Youngchul Sung · 2022
Later among the works it cites.
Agent-time attention for sparse rewards multi-agent reinforcement learning
Jennifer She, Jayesh K Gupta, and Mykel J Kochenderfer · 2022
Later among the works it cites.
The surprising effectiveness of ppo in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu · 2022
Later among the works it cites.
Shaq: Incorporating shapley value theory into multi-agent q-learning
Jianhong Wang, Yuan Zhang, Yunjie Gu, and Tae-Kyun Kim · 2022
Later among the works it cites.
Hossein Haeri, Reza Ahmadzadeh, and Kshitij Jerath · 2022
Later among the works it cites.
D3c: Reducing the price of anarchy in multi-agent learning
Ian Gemp, Kevin R McKee, Richard Everett, Edgar Duéñez-Guzmán, Yoram Bachrach, David Balduzzi, and Andrea Tacchetti · 2022
Later among the works it cites.
Model-Free Opponent Shaping
Chris Lu, Timon Willi, Christian Schroeder de Witt, and Jakob Foerster · 2022
Later among the works it cites.
Towards a standardised performance evaluation protocol for cooperative marl
Rihab Gorsane, Omayma Mahjoub, Ruan John de Kock, Roland Dubb, Siddarth Singh, and Arnu Pretorius · 2022
Later among the works it cites.
A social path to human-like artificial intelligence
Edgar A Duéñez-Guzmán, Suzanne Sadedin, Jane X Wang, Kevin R McKee, and Joel Z Leibo · 2023
Closest in time.
Beyond the matrix: Experimental approaches to studying social-ecological systems
Uri Hertz, Raphael Koster, Marco Janssen, and Joel Z Leibo · 2023
Closest in time.
Peter Sunehag, Alexander Sasha Vezhnevets, Edgar Duéñez-Guzmán, Igor Mordach, and Joel Z Leibo · 2023
Closest in time.
Heterogeneous social value orientation leads to meaningful diversity in sequential social dilemmas
Udari Madhushani, Kevin R McKee, John P Agapiou, Joel Z Leibo, Richard Everett, Thomas Anthony, Edward Hughes, Karl Tuyls, and Edgar A Duéñez-Guzmán · 2023
Closest in time.
Pecan: Leveraging policy ensemble for context-aware zero-shot human-ai coordination
Xingzhou Lou, Jiaxian Guo, Junge Zhang, Jun Wang, Kaiqi Huang, and Yali Du · 2023
Closest in time.
Cooperative open-ended learning framework for zero-shot coordination
Yang Li, Shao Zhang, Jichen Sun, Yali Du, Ying Wen, Xinbing Wang, and Wei Pan · 2023
Closest in time.
Stas: Spatial-temporal return decomposition for multi-agent reinforcement learning
Sirui Chen, Zhaowei Zhang, Yali Du, and Yaodong Yang · 2023
Closest in time.
Resolving social dilemmas with minimal reward transfer
Richard Willis, Yali Du, Joel Z Leibo, and Michael Luck · 2023
Closest in time.
Get it in writing: Formal contracts mitigate social dilemmas in multi-agent rl
Phillip JK Christoffersen, Andreas A Haupt, and Dylan Hadfield-Menell · 2023
Closest in time.
Resolving social dilemmas through reward transfer commitments
Richard Willis and Michael Luck · 2023
Closest in time.
Auto-Aligning Multiagent Incentives with Global Objectives
Minae Kwon, John Agapiou, Edgar Duéñez-Guzmán, Georgios Piliouras, Kalesha Bullard, and Ian Gemp · 2023
Closest in time.
A learning agent that acquires social norms from public sanctions in decentralized multi-agent settings
Eugene Vinitsky, Raphael Köster, John P Agapiou, Edgar A Duéñez-Guzmán, Alexander S Vezhnevets, and Joel Z Leibo · 2023
Closest in time.
Learning to Participate through Trading of Reward Shares
Kyrill Schmid, Michael Kölle, and Tim Matheis · 2023
Closest in time.
Metagpt: Meta programming for multi-agent collaborative framework, 2023
Sirui Hong, Xiawu Zheng, Jonathan Chen, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, and Chenglin Wu · 2023
Closest in time.
Proagent: Building proactive cooperative ai with large language models
Ceyao Zhang, Kaijie Yang, Siyi Hu, Zihao Wang, Guanghe Li, Yihang Sun, Cheng Zhang, Zhaowei Zhang, Anji Liu, Song-Chun Zhu, et al · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior, 2023
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein · 2023
Closest in time.
Alexander Sasha Vezhnevets, John P. Agapiou, Avia Aharon, Ron Ziv, Jayd Matyas, Edgar A. Duéñez Guzmán, William A. Cunningham, Simon Osindero, Danny Karmon, and Joel Z. Leibo · 2023
Closest in time.
Prosocial learning agents solve generalized stag hunts better than selfish ones
Alexander Peysakhovich and Adam Lerer · 2044
Closest in time.