Fetching the paper…
Reading the bibliography…
Multiagent reinforcement learning (MARL) is commonly considered to suffer from non-stationary environments and exponentially increasing policy space.
Learning and executing generalized robot plans
R. Fikes, P. E. Hart, and N. J. Nilsson · 1972
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
M. Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
R. Parr and S. J. Russell · 1997
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. P. Singh · 1999
Earlier work this paper cites.
Hierarchical multi-agent reinforcement learning
R. Makar, S. Mahadevan, and M. Ghavamzadeh · 2001
Earlier work this paper cites.
Learning to communicate and act using hierarchical reinforcement learning
M. Ghavamzadeh and S. Mahadevan · 2004
Earlier work this paper cites.
Multi-agent shared hierarchy reinforcement learning
N. Mehta and P. Tadepalli · 2005
Earlier work this paper cites.
Hierarchical multi-agent reinforcement learning
M. Ghavamzadeh, S. Mahadevan, and R. Makar · 2006
Earlier work this paper cites.
An overview of recent progress in the study of distributed multi-agent coordination
Y. Cao, W. Yu, W. Ren, and G. Chen · 2013
Earlier work this paper cites.
Coordinating multi-agent reinforcement learning with limited communication
C. Zhang and V. R. Lesser · 2013
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
J. N. Foerster, Y. M. Assael, N. Freitas, and S. Whiteson · 2016
Cited alongside, same era.
Hypernetworks
D. Ha, A. M. Dai, and Q. V. Le · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch · 2017
Later among the works it cites.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
S. Omidshafiei, J. Pazis, C. Amato, J. P. How, and J. Vian · 2017
Later among the works it cites.
P. Peng, Y. Wen, Y. Yang, Q. Yuan, Z. Tang, H. Long, and J. Wang · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Cited alongside, same era.
Learning multiagent communication with backpropagation
S. Sukhbaatar, A. Szlam, and R. Fergus · 2016
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
H. v. Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Stabilising experience replay for deep multi-agent reinforcement learning
J. N. Foerster, N. Nardelli, G. Farquhar, T. Afouras, P. H. S. Torr, P. Kohli, and S. Whiteson · 2017
Cited alongside, same era.
#Exploration: A study of count-based exploration for deep reinforcement learning
H. Tang, R. Houthooft, D. Foote, A. Stooke, X. Chen, Y. Duan, J. Schulman, F. D. Turck, and P. Abbeel · 2017
Later among the works it cites.
Counterfactual multi-agent policy gradients
J. N. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson · 2018
Closest in time.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2018
Closest in time.
QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning
T. Rashid, M. Samvelyan, C. Schröder de Witt, G. Farquhar, J. N. Foerster, and S. Whiteson · 2018
Closest in time.