Fetching the paper…
Reading the bibliography…
Multi-agent settings in the real world often involve tasks with varying types and quantities of agents and non-agent entities; however, common patterns of behavior often emerge among these agents/entities.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Claus, C. and Boutilier, C · 1998
Earlier work this paper cites.
Computing factored value functions for policies in structured mdps
Koller, D. and Parr, R · 1999
Earlier work this paper cites.
Distributed Value Functions
Schneider, J., Wong, W.-K., Moore, A., and Riedmiller, M · 1999
Earlier work this paper cites.
Q-decomposition for reinforcement learning agents
Russell, S. and Zimdars, A. L · 2003
Earlier work this paper cites.
Exploiting locality of interaction in factored dec-pomdps
Oliehoek, F. A., Spaan, M. T., Vlassis, N., and Whiteson, S · 2008
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2015
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Hausknecht, M. J. and Stone, P · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Cooperation and Communication in Multiagent Deep Reinforcement Learning
Hausknecht, M. J · 2016
Earlier work this paper cites.
A concise introduction to decentralized POMDPs , volume 1
Oliehoek, F. A., Amato, C., et al · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Hypernetworks
Ha, D., Dai, A. M., and Le, Q. V · 2017
Cited alongside, same era.
A unified Game-Theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Perolat, J., Silver, D., and Graepel, T · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I · 2017
Cited alongside, same era.
The impact of diversity on optimal control policies for heterogeneous robot swarms
Prorok, A., Hsieh, M. A., and Kumar, V · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Actor-attention-critic for multi-agent reinforcement learning
Iqbal, S. and Sha, F · 2019
Later among the works it cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
Lee, J., Lee, Y., Kim, J., Kosiorek, A., Choi, S., and Teh, Y. W · 2019
Later among the works it cites.
Maven: Multi-agent variational exploration
Mahajan, A., Rashid, T., Samvelyan, M., and Whiteson, S · 2019
Later among the works it cites.
The starcraft multi-agent challenge
Samvelyan, M., Rashid, T., Schroeder de Witt, C., Farquhar, G., Nardelli, N., Rudner, T. G., Hung, C.-M., Torr, P. H., Foerster, J., and Whiteson, S · 2019
Later among the works it cites.
Multi-agent common knowledge reinforcement learning
Schroeder de Witt, C., Foerster, J., Farquhar, G., Torr, P., Boehmer, W., and Whiteson, S · 2019
Later among the works it cites.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emergent complexity via multi-agent competition
Bansal, T., Pachocki, J., Sidor, S., Sutskever, I., and Mordatch, I · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Cited alongside, same era.
Learning attentional communication for multi-agent cooperation
Jiang, J. and Lu, Z · 2018
Cited alongside, same era.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., Schroeder, C., Farquhar, G., Foerster, J., and Whiteson, S · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., and Graepel, T · 2018
Cited alongside, same era.
Learning transferable cooperative behavior in multi-agent teams
Agarwal, A., Kumar, S., and Sycara, K · 2019
Cited alongside, same era.
Son, K., Kim, D., Kang, W. J., Hostallero, D. E., and Yi, Y · 2019
Later among the works it cites.
Deep coordination graphs
Böhmer, W., Kurin, V., and Whiteson, S · 2020
Closest in time.
Deep Multi-Agent Reinforcement Learning in Starcraft II
Burden, N · 2020
Closest in time.
Evolutionary population curriculum for scaling multi-agent reinforcement learning
Long, Q., Zhou, Z., Gupta, A., Fang, F., Wu, Y., and Wang, X · 2020
Closest in time.
Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., de Witt, C. S., Farquhar, G., Foerster, J., and Whiteson, S · 2020
Closest in time.
Qatten: A general framework for cooperative multiagent reinforcement learning
Yang, Y., Hao, J., Liao, B., Shao, K., Chen, G., Liu, W., and Tang, H · 2020
Closest in time.
{UPD}et: Universal multi-agent {rl} via policy decoupling with transformers
Hu, S., Zhu, F., Chang, X., and Liang, X · 2021
Closest in time.
{RODE}: Learning roles to decompose multi-agent tasks
Wang, T., Gupta, T., Mahajan, A., Peng, B., Whiteson, S., and Zhang, C · 2021
Closest in time.