Fetching the paper…
Reading the bibliography…
In many real-world settings, a team of agents must coordinate their behaviour while acting in a decentralised way.
Learning from delayed rewards
Watkins, C · 1989
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M · 1993
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Approximation theory of the mlp model in neural networks
Pinkus, A · 1999
Earlier work this paper cites.
Multiagent Planning with Factored MDPs
Guestrin, C., Koller, D., and Parr, R · 2002
Earlier work this paper cites.
Multiagent reinforcement learning for multi-robot systems: A survey
Yang, E. and Gu, D · 2004
Earlier work this paper cites.
Collaborative Multiagent Reinforcement Learning by Payoff Propagation
Kok, J. R. and Vlassis, N · 2006
Earlier work this paper cites.
A Comprehensive Survey of Multiagent Reinforcement Learning
Busoniu, L., Babuska, R., and De Schutter, B · 2008
Earlier work this paper cites.
Optimal and Approximate Q-value Functions for Decentralized POMDPs
Oliehoek, F. A., Spaan, M. T. J., and Vlassis, N · 2008
Earlier work this paper cites.
Incorporating functional knowledge in neural networks
Dugas, C., Bengio, Y., Blisle, F., Nadeau, C., and Garcia, R · 2009
Earlier work this paper cites.
An Overview of Recent Progress in the Study of Distributed Multi-agent Coordination
Cao, Y., Yu, W., Ren, W., and Chen, G · 2012
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2014
Cited alongside, same era.
Deep Recurrent Q-Learning for Partially Observable MDPs
Hausknecht, M. and Stone, P · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., and others · 2015
Cited alongside, same era.
Learning to play guess who? and inventing a grounded language as a consequence
Jorge, E., Kågebäck, M., and Gustavsson, E · 2016
Cited alongside, same era.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
Kraemer, L. and Banerjee, B · 2016
Cited alongside, same era.
A Concise Introduction to Decentralized POMDPs
Oliehoek, F. A. and Amato, C · 2016
Guided Deep Reinforcement Learning for Swarm Systems
Hüttenrauch, M., Šošić, A., and Neumann, G · 2017
Later among the works it cites.
Multi-agent reinforcement learning in sequential social dilemmas
Leibo, J. Z., Zambaldi, V., Lanctot, M., Marecki, J., and Graepel, T · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I · 2017
Later among the works it cites.
Deep Decentralized Multi-task Multi-Agent RL under Partial Observability
Omidshafiei, S., Pazis, J., Amato, C., How, J. P., and Vian, J · 2017
Later among the works it cites.
Peng, P., Wen, Y., Yang, Y., Yuan, Q., Tang, Z., Long, H., and Wang, J · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning multiagent communication with backpropagation
Sukhbaatar, S., Fergus, R., and others · 2016
Cited alongside, same era.
TorchCraft: a Library for Machine Learning Research on Real-Time Strategy Games
Synnaeve, G., Nardelli, N., Auvolat, A., Chintala, S., Lacroix, T., Lin, Z., Richoux, F., and Usunier, N · 2016
Cited alongside, same era.
Stabilising Experience Replay for Deep Multi-Agent Reinforcement Learning
Foerster, J., Nardelli, N., Farquhar, G., Afouras, T., Torr, P. H. S., Kohli, P., and Whiteson, S · 2017
Cited alongside, same era.
Cooperative Multi-agent Control Using Deep Reinforcement Learning
Gupta, J. K., Egorov, M., and Kochenderfer, M · 2017
Cited alongside, same era.
HyperNetworks
Ha, D., Dai, A., and Le, Q. V · 2017
Cited alongside, same era.
Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., and Graepel, T · 2017
Later among the works it cites.
Multiagent cooperation and competition with deep reinforcement learning
Tampuu, A., Matiisen, T., Kodelja, D., Kuzovkin, I., Korjus, K., Aru, J., Aru, J., and Vicente, R · 2017
Later among the works it cites.
Episodic Exploration for Deep Deterministic Policies: An Application to StarCraft Micromanagement Tasks
Usunier, N., Synnaeve, G., Lin, Z., and Chintala, S · 2017
Later among the works it cites.
StarCraft II: A New Challenge for Reinforcement Learning
Vinyals, O., Ewalds, T., Bartunov, S., Georgiev, P., Vezhnevets, A. S., Yeo, M., Makhzani, A., Küttler, H., Agapiou, J., Schrittwieser, J., Quan, J., Gaffney, S., Petersen, S., Simonyan, K., Schaul, T., van Hasselt, H., Silver, D., Lillicrap, T., Calderone, K., Keet, P., Brunasso, A., Lawrence, D., Ekermo, A., Repp, J., and Tsing, R · 2017
Later among the works it cites.
Counterfactual multi-agent policy gradients
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Closest in time.