Fetching the paper…
Reading the bibliography…
The exploitation of extra state information has been an active research area in multi-agent reinforcement learning (MARL).
Probabilistic recursive reasoning for multi-agent reinforcement learning
Wen, Y.; Yang, Y.; Luo, R.; Wang, J.; and Pan, W. 2019 · 1901
Earlier work this paper cites.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Son, K.; Kim, D.; Kang, W. J.; Hostallero, D. E.; and Yi, Y. 2019 · 1905
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016 · 1937
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M. 1993 · 1993
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R.; and Tsitsiklis, J. N. 2000 · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S.; McAllester, D. A.; Singh, S. P.; and Mansour, Y. 2000 · 2000
Earlier work this paper cites.
Counterfactual Multi-Agent Reinforcement Learning with Graph Convolution Communication
Su, J.; Adams, S.; and Beling, P. A. 2020 · 2004
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Bu, L.; Babu, R.; De Schutter, B.; et al. 2008 · 2008
Earlier work this paper cites.
Incorporating Functional Knowledge in Neural Networks
Dugas, C.; Bengio, Y.; Bélisle, F.; Nadeau, C.; and Garcia, R. 2009 · 2009
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J.; Gulcehre, C.; Cho, K.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Hausknecht, M.; and Stone, P. 2015 · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2015 · 2015
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J.; Assael, I. A.; De Freitas, N.; and Whiteson, S. 2016 · 2016
Cited alongside, same era.
Ha, D.; Dai, A.; and Le, Q. V. 2016 · 2016
Cited alongside, same era.
Usunier, N.; Synnaeve, G.; Lin, Z.; and Chintala, S. 2016 · 2016
Cited alongside, same era.
Health-aware hierarchical control for smart manufacturing using reinforcement learning
Choo, B. Y.; Adams, S.; and Beling, P. 2017 · 2017
Counterfactual multi-agent policy gradients
Foerster, J. N.; Farquhar, G.; Afouras, T.; Nardelli, N.; and Whiteson, S. 2018 · 2018
Later among the works it cites.
Learning attentional communication for multi-agent cooperation
Jiang, J.; and Lu, Z. 2018 · 2018
Later among the works it cites.
Credit assignment for collective multiagent RL with global rewards
Nguyen, D. T.; Kumar, A.; and Lau, H. C. 2018 · 2018
Later among the works it cites.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T.; Samvelyan, M.; De Witt, C. S.; Farquhar, G.; Foerster, J.; and Whiteson, S. 2018 · 2018
Later among the works it cites.
Multiagent soft q-learning
Wei, E.; Wicke, D.; Freelan, D.; and Luke, S. 2018 · 2018
Later among the works it cites.
TarMAC: Targeted Multi-Agent Communication
Das, A.; Gervet, T.; Romoff, J.; Batra, D.; Parikh, D.; Rabbat, M.; and Pineau, J. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stabilising experience replay for deep multi-agent reinforcement learning
Foerster, J.; Nardelli, N.; Farquhar, G.; Afouras, T.; Torr, P. H.; Kohli, P.; and Whiteson, S. 2017 · 2017
Cited alongside, same era.
Cooperative multi-agent control using deep reinforcement learning
Gupta, J. K.; Egorov, M.; and Kochenderfer, M. 2017 · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, O. P.; and Mordatch, I. 2017 · 2017
Cited alongside, same era.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games
Peng, P.; Yuan, Q.; Wen, Y.; Yang, Y.; Tang, Z.; Long, H.; and Wang, J. 2017 · 2017
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning
Sunehag, P.; Lever, G.; Gruslys, A.; Czarnecki, W. M.; Zambaldi, V.; Jaderberg, M.; Lanctot, M.; Sonnerat, N.; Leibo, J. Z.; Tuyls, K.; et al. 2017 · 2017
Cited alongside, same era.
Multiagent cooperation and competition with deep reinforcement learning
Tampuu, A.; Matiisen, T.; Kodelja, D.; Kuzovkin, I.; Korjus, K.; Aru, J.; Aru, J.; and Vicente, R. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Interaction-aware Decision Making with Adaptive Strategies under Merging Scenarios
Hu, Y.; Nakhaei, A.; Tomizuka, M.; and Fujimura, K. 2019 · 2019
Later among the works it cites.
Actor-attention-critic for multi-agent reinforcement learning
Iqbal, S.; and Sha, F. 2019 · 2019
Later among the works it cites.
The starcraft multi-agent challenge
Samvelyan, M.; Rashid, T.; Schroeder de Witt, C.; Farquhar, G.; Nardelli, N.; Rudner, T. G.; Hung, C.-M.; Torr, P. H.; Foerster, J.; and Whiteson, S. 2019 · 2019
Later among the works it cites.
Optimal payoff functions for members of collectives
Wolpert, D. H.; and Tumer, K. 2002 · 2019
Later among the works it cites.