Fetching the paper…
Reading the bibliography…
Most recently developed approaches to cooperative multi-agent reinforcement learning in the \emph{centralized training with decentralized execution} setting involve estimating a centralized, joint value function.
The StarCraft Multi-Agent Challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philip H. S. Torr, Jakob Foerster, and Shimon Whiteson · 1902
Earlier work this paper cites.
Truly Proximal Policy Optimization
Yuhui Wang, Hao He, Chao Wen, and Xiaoyang Tan · 1903
Earlier work this paper cites.
QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi · 1905
Earlier work this paper cites.
Exploration with Unreliable Intrinsic Reward in Multi-Agent Reinforcement Learning
Wendelin Böhmer, Tabish Rashid, and Shimon Whiteson · 1906
Earlier work this paper cites.
Wendelin Böhmer, Vitaly Kurin, and Shimon Whiteson · 1910
Earlier work this paper cites.
MAVEN: Multi-Agent Variational Exploration
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, and Shimon Whiteson · 1910
Earlier work this paper cites.
A Complexity Analysis of Cooperative Mechanisms in Reinforcement Learning
Steven Whitehead · 1991
Earlier work this paper cites.
Multi-Agent Reinforcement Learning: Independent vs. Cooperative Agents
Ming Tan · 1993
Earlier work this paper cites.
The Dynamics of Reinforcement Learning in Cooperative Multiagent Systems
C. Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Deep Multi-Agent Reinforcement Learning for Decentralized Continuous Cooperative Control
Christian Schroeder de Witt, Bei Peng, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Böhmer, and Shimon Whiteson · 2003
Earlier work this paper cites.
Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2003
Earlier work this paper cites.
Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2005
Cited alongside, same era.
Cooperative Multi-Agent Learning: The State of the Art
Liviu Panait and Sean Luke · 2005
Cited alongside, same era.
Distributed agent-based air traffic flow management
Kagan Tumer and Adrian Agogino · 2007
Cited alongside, same era.
A Comprehensive Survey of Multiagent Reinforcement Learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2008
Cited alongside, same era.
Revisiting Design Choices in Proximal Policy Optimization
Chloe Ching-Yun Hsu, Celestine Mendler-Dünner, and Moritz Hardt · 2009
Cited alongside, same era.
Lenient Learning in Independent-Learner Stochastic Cooperative Games
Ermo Wei and Sean Luke · 2016
Later among the works it cites.
Counterfactual Multi-Agent Policy Gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2017
Later among the works it cites.
Guided Deep Reinforcement Learning for Swarm Systems
Maximilian Hüttenrauch, Adrian Šošić, and Gerhard Neumann · 2017
Later among the works it cites.
Trust Region Policy Optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yongcan Cao, Wenwu Yu, Wei Ren, and Guanrong Chen · 2012
Cited alongside, same era.
Multiagent Learning: Basics, Challenges, and Prospects
Karl Tuyls and Gerhard Weiss · 2012
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
Landon Kraemer and Bikramjit Banerjee · 2016
Cited alongside, same era.
Asynchronous Methods for Deep Reinforcement Learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
A Concise Introduction to Decentralized POMDPs
Frans A. Oliehoek and Christopher Amato · 2016
Cited alongside, same era.
Value Function Clipping for PPO · Issue #136 · ikostrikov/pytorch-a2c-ppo-acktr-gail
ikostrikov
Cited in the paper.
Later among the works it cites.
Value-Decomposition Networks For Cooperative Multi-Agent Learning
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel · 2017
Later among the works it cites.
StarCraft II: A New Challenge for Reinforcement Learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, John Quan, Stephen Gaffney, Stig Petersen, Karen Simonyan, Tom Schaul, Hado van Hasselt, David Silver, Timothy Lillicrap, Kevin Calderone, Paul Keet, Anthony Brunasso, David Lawrence, Anders Ekermo, Jacob Repp, and Rodney Tsing · 2017
Later among the works it cites.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2018
Later among the works it cites.
Reinforcement Learning, second edition: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms
Anonymous · 2020
Closest in time.