Fetching the paper…
Reading the bibliography…
Centralized Training for Decentralized Execution, where training is done in a centralized offline fashion, has become a popular solution paradigm in Multi-Agent Reinforcement Learning.
The StarCraft Multi-Agent Challenge
Samvelyan, M.; Rashid, T.; de Witt, C. S.; Farquhar, G.; Nardelli, N.; Rudner, T. G. J.; Hung, C.-M.; Torr, P. H. S.; Foerster, J.; and Whiteson, S. 2019 · 1902
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P.; Littman, M. L.; and Cassandra, A. R. 1998 · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning
Sutton, R.; Precup, D.; and Singh, S. 1999 · 1999
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R.; and Tsitsiklis, J. N. 2000 · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S.; McAllester, D. A.; Singh, S. P.; and Mansour, Y. 2000 · 2000
Earlier work this paper cites.
Taming decentralized POMDPs: Towards efficient policy computation for multiagent settings
Nair, R.; Tambe, M.; Yokoo, M.; Pynadath, D.; and Marsella, S. 2003 · 2003
Earlier work this paper cites.
Bounded policy iteration for decentralized POMDPs
Bernstein, D. S.; Hansen, E. A.; and Zilberstein, S. 2005 · 2005
Earlier work this paper cites.
Optimizing Memory-Bounded Controllers for Decentralized POMDPs
Amato, C.; Bernstein, D. S.; and Zilberstein, S. 2007 · 2007
Earlier work this paper cites.
Improved memory-bounded dynamic programming for decentralized POMDPs
Seuken, S.; and Zilberstein, S. 2007 · 2007
Earlier work this paper cites.
Optimal and Approximate Q-value Functions for Decentralized POMDPs
Oliehoek, F. A.; Spaan, M. T. J.; and Vlassis, N. A. 2008 · 2008
Earlier work this paper cites.
Incremental policy generation for finite-horizon DEC-POMDPs
Amato, C.; Dibangoye, J. S.; and Zilberstein, S. 2009 · 2009
Earlier work this paper cites.
Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
Matignon, L.; Laurent, G. J.; and Le Fort-Piat, N. 2012 · 2012
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J.; Assael, I. A.; De Freitas, N.; and Whiteson, S. 2016 · 2016
Cited alongside, same era.
A Concise Introduction to Decentralized POMDPs
Oliehoek, F. A.; and Amato, C. 2016 · 2016
Cited alongside, same era.
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; and Mordatch, I. 2017 · 2017
Cited alongside, same era.
Cooperative multi-agent policy gradient
Bono, G.; Dibangoye, J. S.; Matignon, L.; Pereyron, F.; and Simonin, O. 2018 · 2018
Cited alongside, same era.
Counterfactual Multi-Agent Policy Gradients
Foerster, J.; Farquhar, G.; Afouras, T.; Nardelli, N.; and Whiteson, S. 2018 · 2018
Emergent Tool Use from Multi-Agent Autocurricula
Baker, B.; Kanitscheider, I.; Markov, T.; Wu, Y.; Powell, G.; McGrew, B.; and Mordatch, I. 2020 · 2020
Later among the works it cites.
Shapley Q-Value: A Local Reward Approach to Solve Global Reward Games
Wang, J.; Zhang, Y.; Kim, T.-K.; and Gu, Y. 2020 · 2020
Later among the works it cites.
Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
Zhou, M.; Liu, Z.; Sui, P.; Li, Y.; and Chung, Y. Y. 2020 · 2020
Later among the works it cites.
Unbiased Asymmetric Actor-Critic for Partially Observable Reinforcement Learning
Baisero, A.; and Amato, C. 2021 · 2021
Later among the works it cites.
Learning Correlated Communication Topology in Multi-Agent Reinforcement learning
Du, Y.; Liu, B.; Moens, V.; Liu, Z.; Ren, Z.; Wang, J.; Chen, X.; and Zhang, H. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reinforcement Learning: An Introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Cited alongside, same era.
LIIR: Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning
Du, Y.; Han, L.; Fang, M.; Dai, T.; Liu, J.; and Tao, D. 2019 · 2019
Cited alongside, same era.
Multi Agent Reinforcement Learning Environments Compilation
Jiang, S. 2019 · 2019
Cited alongside, same era.
Multi-agent common knowledge reinforcement learning
Schroeder de Witt, C.; Foerster, J.; Farquhar, G.; Torr, P.; Boehmer, W.; and Whiteson, S. 2019 · 2019
Cited alongside, same era.
R-MADDPG for Partially Observable Environments and Limited Communication
Wang, R. E.; Everett, M.; and How, J. P. 2019 · 2019
Cited alongside, same era.
Deep implicit coordination graphs for multi-agent reinforcement learning
Li, S.; Gupta, J. K.; Morales, P.; Allen, R.; and Kochenderfer, M. J. 2021 · 2021
Later among the works it cites.
Contrasting Centralized and Decentralized Critics in Multi-Agent Reinforcement Learning
Lyu, X.; Xiao, Y.; Daley, B.; and Amato, C. 2021 · 2021
Later among the works it cites.
Multi-Agent Graph-Attention Communication and Teaming
Niu, Y.; Paleja, R.; and Gombolay, M. 2021 · 2021
Later among the works it cites.
Value-Decomposition Multi-Agent Actor-Critics
Su, J.; Adams, S.; and Beling, P. A. 2021 · 2021
Later among the works it cites.
Off-Policy Multi-Agent Decomposed Policy Gradients
Wang, Y.; Han, B.; Wang, T.; Dong, H.; and Zhang, C. 2021 · 2021
Later among the works it cites.