Fetching the paper…
Reading the bibliography…
The necessity for cooperation among intelligent machines has popularised cooperative multi-agent reinforcement learning (MARL) in AI research.
Non-cooperative games
John Nash · 1951
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
A generalized theorem of the maximum
Lawrence M Ausubel and Raymond J Deneckere · 1993
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Dynamic noncooperative game theory
Tamer Başar and Geert Jan Olsder · 1998
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. Mcallester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
Daniel S Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks
Georgios Papoudakis, Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht · 2006
Earlier work this paper cites.
The logit-response dynamics
Carlos Alós-Ferrer and Nick Netzer · 2010
Earlier work this paper cites.
An overview of recent progress in the study of distributed multi-agent coordination
Yongcan Cao, Wenwu Yu, Wei Ren, and Guanrong Chen · 2012
Earlier work this paper cites.
Competitive Markov decision processes
Jerzy Filar and Koos Vrieze · 2012
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
A concise introduction to decentralized POMDPs
Frans A Oliehoek and Christopher Amato · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas · 2016
Cited alongside, same era.
Guided deep reinforcement learning for swarm systems
Maximilian Hüttenrauch, Adrian Šošić, and Gerhard Neumann · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games. arxiv 2017
P Peng, Q Yuan, Y Wen, Y Yang, Z Tang, H Long, and J Wang · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, F. Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Modelling bounded rationality in multi-agent interactions by generalized recursive reasoning
Ying Wen, Yaodong Yang, and Jun Wang · 2020
Later among the works it cites.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yaodong Yang and Jun Wang · 2020
Later among the works it cites.
Multi-agent determinantal q-learning
Yaodong Yang, Ying Wen, Jun Wang, Liheng Chen, Kun Shao, David Mguni, and Weinan Zhang · 2020
Later among the works it cites.
Bi-level actor-critic for multi-agent coordination
Haifeng Zhang, Weizhe Chen, Zeren Huang, Minne Li, Yaodong Yang, Weinan Zhang, and Jun Wang · 2020
Later among the works it cites.
Scaling multi-agent reinforcement learning with selective parameter sharing
Filippos Christianos, Georgios Papoudakis, Muhammad A Rahman, and Stefano V Albrecht · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Heterogeneous multi-agent deep reinforcement learning for traffic lights control
Jeancarlo Arguello Calvo and Ivana Dusparic · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Ian Gemp, Brian McWilliams, Claire Vernade, and Thore Graepel · 2021
Later among the works it cites.
Settling the variance of multi-agent policy gradients
Jakub Grudzien Kuba, Muning Wen, Linghui Meng, Haifeng Zhang, David Mguni, Jun Wang, Yaodong Yang, et al · 2021
Later among the works it cites.
Learning in nonzero-sum stochastic games with potentials
David H Mguni, Yutong Wu, Yali Du, Yaodong Yang, Ziyi Wang, Minne Li, Ying Wen, Joel Jennings, and Jun Wang · 2021
Later among the works it cites.
Facmac: Factored multi-agent centralised policy gradients
Bei Peng, Tabish Rashid, Christian Schroeder de Witt, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Böhmer, and Shimon Whiteson · 2021
Later among the works it cites.
Pettingzoo: Gym for multi-agent reinforcement learning
J Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar, Ananth Hari, Ryan Sullivan, Luis S Santos, Clemens Dieffendahl, Caroline Horsch, Rodrigo Perez-Vicente, et al · 2021
Later among the works it cites.
Rode: Learning roles to decompose multi-agent tasks
Tonghan Wang, Tarun Gupta, Anuj Mahajan, Bei Peng, Shimon Whiteson, and Chongjie Zhang · 2021
Later among the works it cites.
Coordinated proximal policy optimization
Zifan Wu, Chao Yu, Deheng Ye, Junge Zhang, Hankz Hankui Zhuo, et al · 2021
Later among the works it cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar · 2021
Later among the works it cites.
Towards human-level bimanual dexterous manipulation with reinforcement learning
Yuanpei Chen, Yaodong Yang, Tianhao Wu, Shengjie Wang, Xidong Feng, Jiechuan Jiang, Zongqing Lu, Stephen Marcus McAleer, Hao Dong, and Song-Chun Zhu · 2022
Later among the works it cites.
Smacv2: An improved benchmark for cooperative multi-agent reinforcement learning
Benjamin Ellis, Skander Moalla, Mikayel Samvelyan, Mingfei Sun, Anuj Mahajan, Jakob N Foerster, and Shimon Whiteson · 2022
Later among the works it cites.
A game-theoretic approach to multi-agent trust region optimization
Ying Wen, Hui Chen, Yaodong Yang, Minne Li, Zheng Tian, Xu Chen, and Jun Wang · 2022
Later among the works it cites.
The surprising effectiveness of PPO in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu · 2022
Later among the works it cites.
Multiagent trust region policy optimization
Hepeng Li and Haibo He · 2023
Closest in time.
Ray rllib documentation: Multi-agent deep deterministic policy gradient (maddpg)
Ray-Team · 2023
Closest in time.
More centralized training, still decentralized execution: Multi-agent conditional policy factorization
Jiangxing Wang, Deheng Ye, and Zongqing Lu · 2023
Closest in time.
Malib: A parallel framework for population-based multi-agent reinforcement learning
Ming Zhou, Ziyu Wan, Hanjing Wang, Muning Wen, Runzhe Wu, Ying Wen, Yaodong Yang, Yong Yu, Jun Wang, and Weinan Zhang · 2023
Closest in time.