Fetching the paper…
Reading the bibliography…
Centralized Training for Decentralized Execution where agents are trained offline in a centralized fashion and execute online in a decentralized manner, has become a popular approach in Multi-Agent Reinforcement Learning (MARL).
A stochastic approximation method
Robbins, H., and Monro, S. (1951) · 1951
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
Kiefer, J., and Wolfowitz, J. (1952) · 1952
Earlier work this paper cites.
Stability of the gradient process in n-person games
Arrow, K. J., and Hurwicz, L. (1960) · 1960
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Claus, C., and Boutilier, C. (1998) · 1998
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R., and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Learning to cooperate via policy search
Peshkin, L., Kim, K.-E., Meuleau, N., and Kaelbling, L. P. (2000) · 2000
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games
Singh, S. P., Kearns, M. J., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Convergence of gradient dynamics with a variable learning rate
Bowling, M., and Veloso, M. (2001) · 2001
Earlier work this paper cites.
Deep multi-agent reinforcement learning for decentralized continuous cooperative control
de Witt, C. S., Peng, B., Kamienny, P.-A., Torr, P., Böhmer, W., and Whiteson, S. (2020) · 2003
Earlier work this paper cites.
Taming decentralized POMDPs: Towards efficient policy computation for multiagent settings
Nair, R., Tambe, M., Yokoo, M., Pynadath, D., and Marsella, S. (2003) · 2003
Earlier work this paper cites.
Bounded policy iteration for decentralized POMDPs
Bernstein, D. S., Hansen, E. A., and Zilberstein, S. (2005) · 2005
Earlier work this paper cites.
Optimizing memory-bounded controllers for decentralized POMDPs
Amato, C., Bernstein, D. S., and Zilberstein, S. (2007) · 2007
Earlier work this paper cites.
Predicting and preventing coordination problems in cooperative Q-learning systems
Fulda, N., and Ventura, D. (2007) · 2007
Earlier work this paper cites.
Optimal and approximate Q-value functions for decentralized POMDPs
Oliehoek, F. A., Spaan, M. T., and Vlassis, N. (2008) · 2008
Earlier work this paper cites.
Theoretical advantages of lenient Q-learners: An evolutionary game theoretic perspective
Panait, L., Tuyls, K., and Luke, S. (2008) · 2008
Earlier work this paper cites.
Incremental policy generation for finite-horizon DEC-POMDPs
Amato, C., Dibangoye, J. S., and Zilberstein, S. (2009) · 2009
Earlier work this paper cites.
Multi-agent learning with policy prediction
Zhang, C., and Lesser, V. (2010) · 2010
Earlier work this paper cites.
Independent reinforcement learners in cooperative Markov games: a survey regarding coordination problems
Matignon, L., Laurent, G. J., and Le Fort-Piat, N. (2012) · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
The Loss Surfaces of Multilayer Networks
Choromanska, A., Henaff, M., Mathieu, M., Ben Arous, G., and LeCun, Y. (2015) · 2015
Earlier work this paper cites.
Escaping from saddle points — online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y. (2015) · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J., Assael, I. A., De Freitas, N., and Whiteson, S. (2016) · 2016
Earlier work this paper cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Cited alongside, same era.
A Concise Introduction to Decentralized POMDPs
Oliehoek, F. A., and Amato, C. (2016) · 2016
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y. I., Tamar, A., Harb, J., Pieter Abbeel, O., and Mordatch, I. (2017) · 2017
Cited alongside, same era.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
Omidshafiei, S., Pazis, J., Amato, C., How, J. P., and Vian, J. (2017) · 2017
Cited alongside, same era.
Cooperative multi-agent policy gradient
Bono, G., Dibangoye, J. S., Matignon, L., Pereyron, F., and Simonin, O. (2018) · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S. (2018) · 2018
Cited alongside, same era.
Macro-action-based deep multi-agent reinforcement learning
Xiao, Y., Hoffman, J., and Amato, C. (2019) · 2019
Later among the works it cites.
Emergent tool use from multi-agent autocurricula
Baker, B., Kanitscheider, I., Markov, T., Wu, Y., Powell, G., McGrew, B., and Mordatch, I. (2020) · 2020
Later among the works it cites.
Option-critic in cooperative multi-agent systems
Chakravorty, J., Ward, P. N., Roy, J., Chevalier-Boisvert, M., Basu, S., Lupu, A., and Precup, D. (2020) · 2020
Later among the works it cites.
Likelihood quantile networks for coordinating multi-agent reinforcement learning
Lyu, X., and Amato, C. (2020) · 2020
Later among the works it cites.
Weighted QMIX: Expanding monotonic value function factorisation
Rashid, T., Farquhar, G., Peng, B., and Whiteson, S. (2020) · 2020
Later among the works it cites.
Multi-agent actor centralized-critic with communication
Simões, D., Lau, N., and Reis, L. P. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning attentional communication for multi-agent cooperation
Jiang, J., and Lu, Z. (2018) · 2018
Cited alongside, same era.
Emergence of grounded compositional language in multi-agent populations
Mordatch, I., and Abbeel, P. (2018) · 2018
Cited alongside, same era.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., de Witt, C. S., Farquhar, G., Foerster, J., and Whiteson, S. (2018) · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Sutton, R. S., and Barto, A. G. (2018) · 2018
Cited alongside, same era.
TarMAC: Targeted multi-agent communication
Das, A., Gervet, T., Romoff, J., Batra, D., Parikh, D., Rabbat, M., and Pineau, J. (2019) · 2019
Cited alongside, same era.
LIIR: Learning individual intrinsic reward in multi-agent reinforcement learning
Du, Y., Han, L., Fang, M., Dai, T., Liu, J., and Tao, D. (2019) · 2019
Cited alongside, same era.
Later among the works it cites.
Dop: Off-policy multi-agent decomposed policy gradients
Wang, Y., Han, B., Wang, T., Dong, H., and Zhang, C. (2020) · 2020
Later among the works it cites.
A finite-time analysis of two time-scale actor-critic methods
Wu, Y. F., Zhang, W., Xu, P., and Gu, Q. (2020) · 2020
Later among the works it cites.
Learning multi-robot decentralized macro-action-based policies via a centralized Q-net
Xiao, Y., Hoffman, J., Xia, T., and Amato, C. (2020) · 2020
Later among the works it cites.
CM3: Cooperative multi-goal multi-stage multi-agent reinforcement learning
Yang, J., Nakhaei, A., Isele, D., Fujimura, K., and Zha, H. (2020) · 2020
Later among the works it cites.
Learning implicit credit assignment for cooperative multi-agent reinforcement learning
Zhou, M., Liu, Z., Sui, P., Li, Y., and Chung, Y. Y. (2020) · 2020
Later among the works it cites.
Contrasting centralized and decentralized critics in multi-agent reinforcement learning
Lyu, X., Xiao, Y., Daley, B., and Amato, C. (2021) · 2021
Later among the works it cites.
Facmac: Factored multi-agent centralised policy gradients
Peng, B., Rashid, T., Schroeder de Witt, C., Kamienny, P.-A., Torr, P., Böhmer, W., and Whiteson, S. (2021) · 2021
Later among the works it cites.
Value-decomposition multi-agent actor-critics
Su, J., Adams, S., and Beling, P. A. (2021) · 2021
Later among the works it cites.
Off-policy multi-agent decomposed policy gradients
Wang, Y., Han, B., Wang, T., Dong, H., and Zhang, C. (2021) · 2021
Later among the works it cites.
Fop: Factorizing optimal joint policy of maximum-entropy multi-agent reinforcement learning
Zhang, T., Li, Y., Wang, C., Xie, G., and Lu, Z. (2021) · 2021
Later among the works it cites.
Unbiased asymmetric reinforcement learning under partial observability
Baisero, A., and Amato, C. (2022) · 2022
Later among the works it cites.
A deeper understanding of state-based critics in multi-agent reinforcement learning
Lyu, X., Baisero, A., Xiao, Y., and Amato, C. (2022) · 2022
Later among the works it cites.
The surprising effectiveness of PPO in cooperative multi-agent games
Yu, C., Velu, A., Vinitsky, E., Gao, J., Wang, Y., Bayen, A., and Wu, Y. (2022) · 2022
Later among the works it cites.
Towards understanding asynchronous advantage actor-critic: Convergence and linear speedup
Shen, H., Zhang, K., Hong, M., and Chen, T. (2023) · 2023
Later among the works it cites.
Convergence of actor-critic with multi-layer neural networks
Tian, H., Olshevsky, A., and Paschalidis, Y. (2024) · 2024
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., et al. (2018) · 2087
Later among the works it cites.