Fetching the paper…
Reading the bibliography…
We explore value-based solutions for multi-agent reinforcement learning (MARL) tasks in the centralized training with decentralized execution (CTDE) regime popularized recently.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M · 1993
Earlier work this paper cites.
Computing factored value functions for policies in structured mdps
Koller, D. and Parr, R · 1999
Earlier work this paper cites.
Multiagent systems: A survey from a machine learning perspective
Stone, P. and Veloso, M · 2000
Earlier work this paper cites.
Collaborative multiagent reinforcement learning by payoff propagation
Kok, J. R. and Vlassis, N · 2006
Earlier work this paper cites.
Optimal and approximate q-value functions for decentralized pomdps
Oliehoek, F. A., Spaan, M. T., and Vlassis, N · 2008
Earlier work this paper cites.
Exploiting structure and agent-centric rewards to promote coordination in large multiagent systems
HolmesParker, C., Taylor, M., Zhan, Y., and Tumer, K · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J., Assael, I. A., de Freitas, N., and Whiteson, S · 2016
Earlier work this paper cites.
A concise introduction to decentralized POMDPs , volume 1
Oliehoek, F. A., Amato, C., et al · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Cited alongside, same era.
Learning multiagent communication with backpropagation
Sukhbaatar, S., Fergus, R., et al · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Cited alongside, same era.
Lenient learning in independent-learner stochastic cooperative games
Wei, E. and Luke, S · 2016
Cited alongside, same era.
Cooperative multi-agent control using deep reinforcement learning
Gupta, J. K., Egorov, M., and Kochenderfer, M · 2017
Multiagent cooperation and competition with deep reinforcement learning
Tampuu, A., Matiisen, T., Kodelja, D., Kuzovkin, I., Korjus, K., Aru, J., Aru, J., and Vicente, R · 2017
Later among the works it cites.
Factorized q-learning for large-scale multi-agent systems
Chen, Y., Zhou, M., Wen, Y., Yang, Y., Su, Y., Zhang, W., Zhang, D., Wang, J., and Liu, H · 2018
Later among the works it cites.
Counterfactual multi-agent policy gradients
Foerster, J. N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Later among the works it cites.
Learning attentional communication for multi-agent cooperation
Jiang, J. and Lu, Z · 2018
Later among the works it cites.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., Schroeder, C., Farquhar, G., Foerster, J., and Whiteson, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., WU, Y., Tamar, A., Harb, J., Pieter Abbeel, O., and Mordatch, I · 2017
Cited alongside, same era.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
Omidshafiei, S., Pazis, J., Amato, C., How, J. P., and Vian, J · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Multiagent planning with factored mdps
Guestrin, C., Koller, D., and Parr, R
Cited in the paper.
Coordinated reinforcement learning
Guestrin, C., Lagoudakis, M., and Parr, R
Cited in the paper.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V. F., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., and Graepel, T · 2018
Later among the works it cites.
Wei, E., Wicke, D., Freelan, D., and Luke, S · 2018
Later among the works it cites.
Mean field multi-agent reinforcement learning
Yang, Y., Luo, R., Li, M., Zhou, M., Zhang, W., and Wang, J · 2018
Later among the works it cites.
Learning to schedule communication in multi-agent reinforcement learning
Kim, D., Moon, S., Hostallero, D., Kang, W. J., Lee, T., Son, K., and Yi, Y · 2019
Closest in time.