Fetching the paper…
Reading the bibliography…
In cooperative multi-agent reinforcement learning (MARL), combining value decomposition with actor-critic enables agents to learn stochastic policies, which are more suitable for the partially observable environment.
Planning, learning and coordination in multiagent decision processes
Craig Boutilier · 1996
Earlier work this paper cites.
Topology , volume 2
James R Munkres · 2000
Earlier work this paper cites.
The Evidential Foundations of Probabilistic Reasoning
David A Schum · 2001
Earlier work this paper cites.
Coordinated reinforcement learning
Carlos Guestrin, Michail Lagoudakis, and Ronald Parr · 2002
Earlier work this paper cites.
Off-policy multi-agent decomposed policy gradients
Yihan Wang, Beining Han, Tonghan Wang, Heng Dong, and Chongjie Zhang · 2007
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
Landon Kraemer and Bikramjit Banerjee · 2016
Earlier work this paper cites.
Cooperative multi-agent control using deep reinforcement learning
Jayesh K Gupta, Maxim Egorov, and Mykel Kochenderfer · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Weighted qmix: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Gregory Farquhar, Bei Peng, and Shimon Whiteson · 2020
Later among the works it cites.
Difference rewards policy gradients
Jacopo Castellini, Sam Devlin, Frans A Oliehoek, and Rahul Savani · 2021
Later among the works it cites.
Trust region policy optimisation in multi-agent reinforcement learning
Jakub Grudzien Kuba, Ruiqing Chen, Munning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang · 2021
Later among the works it cites.
Value-decomposition multi-agent actor-critics
Jianyu Su, Stephen Adams, and Peter A Beling · 2021
Later among the works it cites.
The surprising effectiveness of ppo in cooperative, multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre Bayen, and Yi Wu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dimitri Bertsekas · 2019
Cited alongside, same era.
Coordinated exploration via intrinsic rewards for multi-agent reinforcement learning
Shariq Iqbal and Fei Sha · 2019
Cited alongside, same era.
Maven: Multi-agent variational exploration
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, and Shimon Whiteson · 2019
Cited alongside, same era.
The StarCraft Multi-Agent Challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philiph H. S. Torr, Jakob Foerster, and Shimon Whiteson · 2019
Cited alongside, same era.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi · 2019
Cited alongside, same era.
Learning transferable cooperative behavior in multi-agent teams
Akshat Agarwal, Sumit Kumar, Katia Sycara, and Michael Lewis · 2020
Cited alongside, same era.
Deep coordination graphs
Wendelin Böhmer, Vitaly Kurin, and Shimon Whiteson · 2020
Cited alongside, same era.
Fop: Factorizing optimal joint policy of maximum-entropy multi-agent reinforcement learning
Tianhao Zhang, Yueheng Li, Chen Wang, Guangming Xie, and Zongqing Lu · 2021
Later among the works it cites.
Episodic multi-agent reinforcement learning with curiosity-driven exploration
Lulu Zheng, Jiarui Chen, Jianhao Wang, Jiamin He, Yujing Hu, Yingfeng Chen, Changjie Fan, Yang Gao, and Chongjie Zhang · 2021
Later among the works it cites.
Multi-agent reinforcement learning for online scheduling in smart factories
Tong Zhou, Dunbing Tang, Haihua Zhu, and Zequn Zhang · 2021
Later among the works it cites.
Revisiting some common practices in cooperative multi-agent reinforcement learning
Wei Fu, Chao Yu, Zelai Xu, Jiaqi Yang, and Yi Wu · 2022
Closest in time.
Difference advantage estimation for multi-agent policy gradients
Yueheng Li, Guangming Xie, and Zongqing Lu · 2022
Closest in time.
Divergence-regularized multi-agent actor-critic
Kefan Su and Zongqing Lu · 2022
Closest in time.
Self-organized polynomial-time coordination graphs
Qianlan Yang, Weijun Dong, Zhizhou Ren, Jianhao Wang, Tonghan Wang, and Chongjie Zhang · 2022
Closest in time.