Fetching the paper…
Reading the bibliography…
Many complex multi-agent systems such as robot swarms control and autonomous vehicle coordination can be modeled as Multi-Agent Reinforcement Learning (MARL) tasks.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Incremental multi-step q-learning
Jing Peng and Ronald J Williams · 1994
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh · 2000
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas · 2003
Earlier work this paper cites.
Value-decomposition multi-agent actor-critics
Jianyu Su, Stephen Adams, and Peter A Beling · 2007
Earlier work this paper cites.
Value-Decomposition Multi-Agent Actor-Critics
Jianyu Su, Stephen Adams, and Peter A. Beling · 2007
Earlier work this paper cites.
Off-Policy Multi-Agent Decomposed Policy Gradients
Yihan Wang, Beining Han, Tonghan Wang, Heng Dong, and Chongjie Zhang · 2007
Earlier work this paper cites.
Learning Implicit Credit Assignment for Multi-Agent Actor-Critic
Meng Zhou, Ziyu Liu, Pengwei Sui, Yixuan Li, and Yuk Ying Chung · 2007
Earlier work this paper cites.
QPLEX: Duplex Dueling Multi-Agent Q-Learning
Jianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu, and Chongjie Zhang · 2008
Earlier work this paper cites.
Pomdps for robotic tasks with mixed observability
Sylvie CW Ong, Shao Wei Png, David Hsu, and Wee Sun Lee · 2009
Earlier work this paper cites.
Coordinated multi-agent reinforcement learning in networked distributed pomdps
Chongjie Zhang and Victor R. Lesser · 2011
Earlier work this paper cites.
An overview of recent progress in the study of distributed multi-agent coordination
Yongcan Cao, Wenwu Yu, Wei Ren, and Guanrong Chen · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
Landon Kraemer and Bikramjit Banerjee · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Guided deep reinforcement learning for swarm systems
Maximilian Hüttenrauch, Adrian Šošić, and Gerhard Neumann · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Multi-agent manipulation via locomotion using hierarchical sim2real
Ofir Nachum, Michael Ahn, Hugo Ponte, Shixiang Gu, and Vikash Kumar · 2019
Later among the works it cites.
The StarCraft Multi-Agent Challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philip H. S. Torr, Jakob Foerster, and Shimon Whiteson · 2019
Later among the works it cites.
QTRAN: learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Hostallero, and Yung Yi · 2019
Later among the works it cites.
What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk, Manu Orsini, Sertan Girgin, Raphael Marinier, Léonard Hussenot, Matthieu Geist, Olivier Pietquin, Marcin Michalski, Sylvain Gelly, and Olivier Bachem · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Value-Decomposition Networks For Cooperative Multi-Agent Learning
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel · 2017
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Jakob N. Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Cited alongside, same era.
QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schröder de Witt, Gregory Farquhar, Jakob N. Foerster, and Shimon Whiteson · 2018
Cited alongside, same era.
Accelerated methods for deep reinforcement learning
Adam Stooke and Pieter Abbeel · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Ermo Wei, Drew Wicke, David Freelan, and Sean Luke · 2018
Cited alongside, same era.
Deep coordination graphs
Wendelin Boehmer, Vitaly Kurin, and Shimon Whiteson · 2020
Later among the works it cites.
Karl Cobbe, Jacob Hilton, Oleg Klimov, and John Schulman · 2020
Later among the works it cites.
Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2020
Later among the works it cites.
Facmac: Factored multi-agent centralised policy gradients
Bei Peng, Tabish Rashid, Christian A Schroeder de Witt, Pierre-Alexandre Kamienny, Philip HS Torr, Wendelin Böhmer, and Shimon Whiteson · 2020
Later among the works it cites.
Weighted QMIX: Expanding Monotonic Value Function Factorisation
Tabish Rashid, Gregory Farquhar, Bei Peng, and Shimon Whiteson · 2020
Later among the works it cites.
QTRAN++: Improved Value Transformation for Cooperative Multi-Agent Reinforcement Learning
Kyunghwan Son, Sungsoo Ahn, Roben Delos Reyes, Jinwoo Shin, and Yung Yi · 2020
Later among the works it cites.
Macro-action-based deep multi-agent reinforcement learning
Yuchen Xiao, Joshua Hoffman, and Christopher Amato · 2020
Later among the works it cites.
Qatten: A General Framework for Cooperative Multiagent Reinforcement Learning
Yaodong Yang, Jianye Hao, Ben Liao, Kun Shao, Guangyong Chen, Wulong Liu, and Hongyao Tang · 2020
Later among the works it cites.
Revisiting peng’s q ( λ \lambda ) for modern reinforcement learning
Tadashi Kozuno, Yunhao Tang, Mark Rowland, Rémi Munos, Steven Kapturowski, Will Dabney, Michal Valko, and David Abel · 2021
Closest in time.