Fetching the paper…
Reading the bibliography…
Centralized Training with Decentralized Execution (CTDE) has recently emerged as a popular framework for cooperative Multi-Agent Reinforcement Learning (MARL), where agents can use additional global state information to guide training in a centralized way and make their own decisions only based on decentralized local policies.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
CTDS: centralized teacher with decentralized student for multi-agent reinforcement learning
Jian Zhao, Xunhan Hu, Mingyu Yang, Wengang Zhou, Jiangcheng Zhu, and Houqiang Li · 1998
Earlier work this paper cites.
Introduction to cognition and communication
Keith Stenning, Jo Calder, and Alex Lascarides · 2006
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, et al · 2015
Earlier work this paper cites.
Learning multiagent communication with backpropagation
Sainbayar Sukhbaatar, Rob Fergus, et al · 2016
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Counterfactual multi-agent policy gradients
Jakob N. Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Earlier work this paper cites.
Learning attentional communication for multi-agent cooperation
Jiechuan Jiang and Zongqing Lu · 2018
Earlier work this paper cites.
QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schröder de Witt, Gregory Farquhar, Jakob N. Foerster, and Shimon Whiteson · 2018
Earlier work this paper cites.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, et al · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Tarmac: Targeted multi-agent communication
Abhishek Das, Théophile Gervet, Joshua Romoff, Dhruv Batra, Devi Parikh, Mike Rabbat, and Joelle Pineau · 2019
Earlier work this paper cites.
Actor-attention-critic for multi-agent reinforcement learning
Shariq Iqbal and Fei Sha · 2019
Earlier work this paper cites.
The starcraft multi-agent challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schröder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philip H. S. Torr, Jakob N. Foerster, and Shimon Whiteson · 2019
Earlier work this paper cites.
QTRAN: learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Hostallero, and Yung Yi · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, et al · 2019
Cited alongside, same era.
Learning individually inferred communication for multi-agent cooperation
Ziluo Ding, Tiejun Huang, and Zongqing Lu · 2020
Cited alongside, same era.
Google research football: A novel reinforcement learning environment
Karol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac, Olivier Bachem, Lasse Espeholt, Carlos Riquelme, Damien Vincent, Marcin Michalski, Olivier Bousquet, and Sylvain Gelly · 2020
Cited alongside, same era.
Learning agent communication under limited bandwidth by message pruning
Hangyu Mao, Zhengchao Zhang, Zhen Xiao, Zhibo Gong, and Yan Ni · 2020
Cited alongside, same era.
Ma2ql: A minimalist approach to fully decentralized multi-agent reinforcement learning
Kefan Su, Siyuan Zhou, Chuang Gan, Xiangjun Wang, and Zongqing Lu · 2022
Later among the works it cites.
The surprising effectiveness of ppo in cooperative, multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre Bayen, and Yi Wu · 2022
Later among the works it cites.
Multi-agent incentive communication via decentralized teammate modeling
Lei Yuan, Jianhao Wang, Fuxiang Zhang, Chenghe Wang, Zongzhang Zhang, Yang Yu, and Chongjie Zhang · 2022
Later among the works it cites.
Smacv2: An improved benchmark for cooperative multi-agent reinforcement learning
Benjamin Ellis, Skander Moalla, Mikayel Samvelyan, Mingfei Sun, Anuj Mahajan, Jakob N Foerster, and Shimon Whiteson · 2023
Closest in time.
Ace: Cooperative multi-agent q-learning with bidirectional action-dependency
Chuming Li, Jie Liu, Yinmin Zhang, Yuhong Wei, Yazhe Niu, Yaodong Yang, Yu Liu, and Wanli Ouyang · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning multi-agent communication with double attentional deep reinforcement learning
Hangyu Mao, Zhengchao Zhang, Zhen Xiao, Zhibo Gong, and Yan Ni · 2020
Cited alongside, same era.
Weighted QMIX: expanding monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Gregory Farquhar, Bei Peng, and Shimon Whiteson · 2020
Cited alongside, same era.
Learning efficient multi-agent communication: An information bottleneck approach
Rundong Wang, Xu He, Runsheng Yu, Wei Qiu, Bo An, and Zinovi Rabinovich · 2020
Cited alongside, same era.
Learning efficient multi-agent communication: An information bottleneck approach
Rundong Wang, Xu He, Runsheng Yu, Wei Qiu, Bo An, and Zinovi Rabinovich · 2020
Cited alongside, same era.
Learning nearly decomposable value functions via communication minimization
Tonghan Wang, Jianhao Wang, Chongyi Zheng, and Chongjie Zhang · 2020
Cited alongside, same era.
Multi-agent deep reinforcement learning for urban traffic light control in vehicular networks
Tong Wu, Pan Zhou, Kai Liu, Yali Yuan, Xiumin Wang, Huawei Huang, and Dapeng Oliver Wu · 2020
Cited alongside, same era.
Action candidate based clipped double q-learning for discrete and continuous action tasks
Haobo Jiang, Jin Xie, and Jian Yang · 2021
Cited alongside, same era.
Contrastive identity-aware learning for multi-agent value decomposition
Shunyu Liu, Yihe Zhou, Jie Song, Tongya Zheng, Kaixuan Chen, Tongtian Zhu, Zunlei Feng, and Mingli Song · 2023
Closest in time.
Scalable multi-agent reinforcement learning through intelligent information aggregation, 2023
Siddharth Nayak, Kenneth Choi, Wenqi Ding, Sydney Dolan, Karthik Gopalakrishnan, and Hamsa Balakrishnan · 2023
Closest in time.
Attention-based recurrence for multi-agent reinforcement learning under stochastic partial observability
Thomy Phan, Fabian Ritz, Philipp Altmann, Maximilian Zorn, Jonas Nüßlein, Michael Kölle, Thomas Gabor, and Claudia Linnhoff-Popien · 2023
Closest in time.
More centralized training, still decentralized execution: Multi-agent conditional policy factorization
Jiangxing Wang, Deheng Ye, and Zongqing Lu · 2023
Closest in time.
PTDE: personalized training with distillated execution for multi-agent reinforcement learning
Yiqun Chen, Hangyu Mao, Tianle Zhang, Shiguang Wu, Bin Zhang, Jianye Hao, Dong Li, Bin Wang, and Hongxing Chang · 2024
Closest in time.
Learning multi-agent communication from graph modeling perspective
Shengchao Hu, Li Shen, Ya Zhang, and Dacheng Tao · 2024
Closest in time.
A black-box approach for non-stationary multi-agent reinforcement learning
Haozhe Jiang, Qiwen Cui, Zhihan Xiong, Maryam Fazel, and Simon S. Du · 2024
Closest in time.
Assigning credit with partial reward decoupling in multi-agent proximal policy optimization, 2024
Aditya Kapoor, Benjamin Freed, Howie Choset, and Jeff Schneider · 2024
Closest in time.
Interaction pattern disentangling for multi-agent reinforcement learning
Shunyu Liu, Jie Song, Yihe Zhou, Na Yu, Kaixuan Chen, Zunlei Feng, and Mingli Song · 2024
Closest in time.
Learning to communicate using contrastive learning
Yat Long Lo, Biswa Sengupta, Jakob Foerster, and Michael Noukhovitch · 2024
Closest in time.
A2PO: towards effective offline reinforcement learning from an advantage-aware perspective
Yunpeng Qing, Shunyu Liu, Jingyuan Cong, Kaixuan Chen, Yihe Zhou, and Mingli Song · 2024
Closest in time.
Temporal prototype-aware learning for active voltage control on power distribution networks
Feiyang Xu, Shunyu Liu, Yunpeng Qing, Yihe Zhou, Yuwen Wang, and Mingli Song · 2024
Closest in time.