Fetching the paper…
Reading the bibliography…
Cooperative multi-agent reinforcement learning is a powerful tool to solve many real-world cooperative tasks, but restrictions of real-world applications may require training the agents in a fully decentralized manner.
The starcraft multi-agent challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philiph H. S. Torr, Jakob Foerster, and Shimon Whiteson · 1902
Earlier work this paper cites.
The starcraft multi-agent challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim GJ Rudner, Chia-Man Hung, Philip HS Torr, Jakob Foerster, and Shimon Whiteson · 1902
Earlier work this paper cites.
The complexity of markov decision processes
Christos H Papadimitriou and John N Tsitsiklis · 1987
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
Martin Lauer and Martin Riedmiller · 2000
Earlier work this paper cites.
Complexity of finite-horizon markov decision process problems
Martin Mundhenk, Judy Goldsmith, Christopher Lusena, and Eric Allender · 2000
Earlier work this paper cites.
Hysteretic q-learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams
Laëtitia Matignon, Guillaume J Laurent, and Nadine Le Fort-Piat · 2007
Earlier work this paper cites.
A survey of ppad-completeness for computing nash equilibria
Paul W Goldberg · 2011
Earlier work this paper cites.
On the computational complexity of stochastic controller optimization in pomdps
Nikos Vlassis, Michael L Littman, and David Barber · 2012
Earlier work this paper cites.
Coordinating multi-agent reinforcement learning with limited communication
Chongjie Zhang and Victor Lesser · 2013
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Decentralized q-learning for stochastic teams and games
Gürdal Arslan and Serdar Yüksel · 2016
Earlier work this paper cites.
Stabilising experience replay for deep multi-agent reinforcement learning
Jakob Foerster, Nantas Nardelli, Gregory Farquhar, Triantafyllos Afouras, Philip HS Torr, Pushmeet Kohli, and Shimon Whiteson · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Earlier work this paper cites.
Lenient multi-agent deep reinforcement learning
Gregory Palmer, Karl Tuyls, Daan Bloembergen, and Rahul Savani · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder De Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Cited alongside, same era.
Fully decentralized multi-agent reinforcement learning with networked agents
Kaiqing Zhang, Zhuoran Yang, Han Liu, Tong Zhang, and Tamer Basar · 2018
Cited alongside, same era.
Actor-attention-critic for multi-agent reinforcement learning
Shariq Iqbal and Fei Sha · 2019
Cited alongside, same era.
Facmac: Factored multi-agent centralised policy gradients
Bei Peng, Tabish Rashid, Christian Schroeder de Witt, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Böhmer, and Shimon Whiteson · 2021
Later among the works it cites.
Hierarchically and cooperatively learning traffic signal control
Bingyu Xu, Yaowei Wang, Zhaozhi Wang, Huizhu Jia, and Zongqing Lu · 2021
Later among the works it cites.
Decentralized learning for optimality in stochastic dynamic teams and games with local control and global state information
Bora Yongacoglu, Gürdal Arslan, and Serdar Yüksel · 2021
Later among the works it cites.
Decentralized policy gradient for nash equilibria learning of general-sum stochastic games
Yan Chen and Tao Li · 2022
Later among the works it cites.
S Rasoul Etesami · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi · 2019
Cited alongside, same era.
Deep coordination graphs
Wendelin Böhmer, Vitaly Kurin, and Shimon Whiteson · 2020
Cited alongside, same era.
Is independent learning all you need in the starcraft multi-agent challenge?
Christian Schroeder de Witt, Tarun Gupta, Denys Makoviichuk, Viktor Makoviychuk, Philip HS Torr, Mingfei Sun, and Shimon Whiteson · 2020
Cited alongside, same era.
Google research football: A novel reinforcement learning environment
Karol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajkac, Olivier Bachem, Lasse Espeholt, Carlos Riquelme, Damien Vincent, Marcin Michalski, Olivier Bousquet, et al · 2020
Cited alongside, same era.
Multi-agent trust region policy optimization
Hepeng Li and Haibo He · 2020
Cited alongside, same era.
Weighted qmix: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Gregory Farquhar, Bei Peng, and Shimon Whiteson · 2020
Cited alongside, same era.
Dop: Off-policy multi-agent decomposed policy gradients
Yihan Wang, Beining Han, Tonghan Wang, Heng Dong, and Chongjie Zhang · 2020
Cited alongside, same era.
Later among the works it cites.
On the convergence of policy gradient methods to nash equilibria in general stochastic games
Angeliki Giannou, Kyriakos Lotidis, Panayotis Mertikopoulos, and Emmanouil-Vasileios Vlatakis-Gkaragkounis · 2022
Later among the works it cites.
I2q: A fully decentralized q-learning algorithm
Jiechuan Jiang and Zongqing Lu · 2022
Later among the works it cites.
Difference advantage estimation for multi-agent policy gradients
Yueheng Li, Guangming Xie, and Zongqing Lu · 2022
Later among the works it cites.
On improving model-free algorithms for decentralized multi-agent reinforcement learning
Weichao Mao, Lin Yang, Kaiqing Zhang, and Tamer Basar · 2022
Later among the works it cites.
Ma2ql: A minimalist approach to fully decentralized multi-agent reinforcement learning
Kefan Su, Siyuan Zhou, Chuang Gan, Xiangjun Wang, and Zongqing Lu · 2022
Later among the works it cites.
The surprising effectiveness of ppo in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu · 2022
Later among the works it cites.
Breaking the curse of multiagents in a large state space: Rl in markov games with independent linear function approximation
Qiwen Cui, Kaiqing Zhang, and Simon Du · 2023
Later among the works it cites.
Learning multi-agent intention-aware communication for optimal multi-order execution in finance
Yuchen Fang, Zhenggang Tang, Kan Ren, Weiqing Liu, Li Zhao, Jiang Bian, Dongsheng Li, Weinan Zhang, Yong Yu, and Tie-Yan Liu · 2023
Later among the works it cites.
Provably efficient reinforcement learning in decentralized general-sum markov games
Weichao Mao and Tamer Başar · 2023
Later among the works it cites.
Multi-agent deep reinforcement learning for multi-robot applications: a survey
James Orr and Ayan Dutta · 2023
Later among the works it cites.
Asynchronous decentralized q-learning: Two timescale analysis by persistence
Bora Yongacoglu, Gürdal Arslan, and Serdar Yüksel · 2023
Later among the works it cites.
f f -divergence policy optimization in fully decentralized cooperative marl, 2024
Kefan Su and Zongqing Lu · 2024
Closest in time.