Fetching the paper…
Reading the bibliography…
This paper introduces the deep coordination graph (DCG) for collaborative multi-agent reinforcement learning.
Exploration with unreliable intrinsic reward in multi-agent reinforcement learning
Böhmer, W., Rashid, T., and Whiteson, S · 1906
Earlier work this paper cites.
A review of cooperative multi-agent deep reinforcement learning
Oroojlooy jadid, A. and Hajinezhad, D · 1908
Earlier work this paper cites.
Multi-agent game abstraction via graph attention neural network, 2019
Liu, Y., Wang, W., Hu, Y., Hao, J., Chen, X., and Gao, Y · 1911
Earlier work this paper cites.
Applying max-sum to teams of mobile sensing agents
Yedidsion, H., Zivan, R., and Farinelli, A · 1976
Earlier work this paper cites.
Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference
Pearl, J · 1988
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. and Dayan, P · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M · 1993
Earlier work this paper cites.
Computing factored value functions for policies in structured mdps
Koller, D. and Parr, R · 1999
Earlier work this paper cites.
Loopy belief propagation for approximate inference: An empirical study
Murphy, K. P., Weiss, Y., and Jordan, M. I · 1999
Earlier work this paper cites.
Loopy belief propagation as a basis for communication in sensor networks
Crick, C. and Pfeffer, A · 2002
Earlier work this paper cites.
Understanding belief propagation and its generalizations
Yedidia, J. S., Freeman, W. T., and Weiss, Y · 2003
Earlier work this paper cites.
Tree consistency and bounds on the performance of the max-product algorithm and its generalizations
Wainwright, M., Jaakkola, T., and Willsky, A · 2004
Earlier work this paper cites.
Collaborative multiagent reinforcement learning by payoff propagation
Kok, J. R. and Vlassis, N · 2006
Earlier work this paper cites.
Biasing coevolutionary search for optimal multiagent behaviors
Panait, L., Luke, S., and Wiegand, R. P · 2006
Earlier work this paper cites.
Bounded approximate decentralised coordination via the max-sum algorithm
Rogers, A. C., Farinelli, A., Stranders, R., and Jennings, N. R · 2011
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2014
Cited alongside, same era.
Integrated power and natural gas model for energy adequacy in short-term operation
Correa-Posada, C. M. and Sánchez-Martin, P · 2015
Cited alongside, same era.
Deep recurrent q-learning for partially observable mdps
Hausknecht, M. J. and Stone, P · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J., Assael, Y., de Freitas, N., and Whiteson, S · 2016
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Later among the works it cites.
QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., de Witt, C. S., Farquhar, G., Foerster, J. N., and Whiteson, S · 2018
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., and Graepel, T · 2018
Later among the works it cites.
Multiagent soft q-learning
Wei, E., Wicke, D., Freelan, D., and Luke, S · 2018
Later among the works it cites.
Mean field multi-agent reinforcement learning
Yang, Y., Luo, R., Li, M., Zhou, M., Zhang, W., and Wang, J · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-agent reinforcement learning as a rehearsal for decentralized planning
Kraemer, L. and Banerjee, B · 2016
Cited alongside, same era.
A concise introduction to decentralized POMDPs
Oliehoek, F. A. and Amato, C · 2016
Cited alongside, same era.
Coordinated deep reinforcement learners for traffic light control
Van der Pol, E. and Oliehoek, F. A · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Advanced control in factory automation: a survey
Dotoli, M., Fay, A., Miśkowicz, M., and Seatzu, C · 2017
Cited alongside, same era.
Stabilising experience replay for deep multi-agent reinforcement learning
Foerster, J., Nardelli, N., Farquhar, G., Torr, P., Kohli, P., and Whiteson, S · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., WU, Y., Tamar, A., Harb, J., Pieter Abbeel, O., and Mordatch, I · 2017
Cited alongside, same era.
Learning from demonstration in the wild
Behbahani, F., Shiarlis, K., Chen, X., Kurin, V., Kasewa, S., Stirbu, C., Gomes, J., Paul, S., Oliehoek, F. A., Messias, J., and Whiteson, S · 2019
Closest in time.
The representational capacity of action-value networks for multi-agent reinforcement learning
Castellini, J., Oliehoek, F. A., Savani, R., and Whiteson, S · 2019
Closest in time.
Multi-agent collaborative exploration through graph-based deep reinforcement learning
Luo, T., Subagdja, B., Wang, D., and Tan, A · 2019
Closest in time.
Magnet: Multi-agent graph network for deep multi-agent reinforcement learning
Malysheva, A., Kudenko, D., and Shpilman, A · 2019
Closest in time.
The StarCraft Multi-Agent Challenge
Samvelyan, M., Rashid, T., de Witt, C. S., Farquhar, G., Nardelli, N., Rudner, T. G. J., Hung, C.-M., Torr, P. H. S., Foerster, J., and Whiteson, S · 2019
Closest in time.
Multi-agent common knowledge reinforcement learning
Schröder de Witt, C. A., Förster, J. N., Farquhar, G., Torr, P. H., Böhmer, W., and Whiteson, S · 2019
Closest in time.
QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Son, K., Kim, D., Kang, W. J., Hostallero, D. E., and Yi, Y · 2019
Closest in time.
Relational forward models for multi-agent learning
Tacchetti, A., Song, H. F., Mediano, P. A. M., Zambaldi, V., Kramár, J., Rabinowitz, N. C., Graepel, T., Botvinick, M., and Battaglia, P. W · 2019
Closest in time.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gulcehre, C., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D · 2019
Closest in time.
Graph convolutional reinforcement learning
Jiang, J., Dun, C., Huang, T., and Lu, Z · 2020
Closest in time.