Fetching the paper…
Reading the bibliography…
Over these years, multi-agent reinforcement learning has achieved remarkable performance in multi-agent planning and scheduling tasks.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M. 1993 · 1993
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G. 1994 · 1994
Earlier work this paper cites.
Taming decentralized POMDPs: Towards efficient policy computation for multiagent settings
Nair, R.; Tambe, M.; Yokoo, M.; Pynadath, D.; and Marsella, S. 2003 · 2003
Earlier work this paper cites.
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
Lowe, R.; WU, Y.; Tamar, A.; Harb, J.; Pieter Abbeel, O.; and Mordatch, I. 2017 · 2017
Earlier work this paper cites.
Heterogeneous Multi-Agent Deep Reinforcement Learning for Traffic Lights Control
Calvo, J. A.; and Dusparic, I. 2018 · 2018
Earlier work this paper cites.
Counterfactual multi-agent policy gradients
Foerster, J.; Farquhar, G.; Afouras, T.; Nardelli, N.; and Whiteson, S. 2018 · 2018
Earlier work this paper cites.
Efficient large-scale fleet management via multi-agent deep reinforcement learning
Lin, K.; Zhao, R.; Xu, Z.; and Zhou, J. 2018 · 2018
Earlier work this paper cites.
Learning robust options
Mankowitz, D.; Mann, T.; Bacon, P.-L.; Precup, D.; and Mannor, S. 2018 · 2018
Earlier work this paper cites.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T.; Samvelyan, M.; Schroeder, C.; Farquhar, G.; Foerster, J.; and Whiteson, S. 2018 · 2018
Earlier work this paper cites.
Simplified Action Decoder for Deep Multi-Agent Reinforcement Learning
Hu, H.; and Foerster, J. N. 2019 · 2019
Earlier work this paper cites.
Disjoint splitting for multi-agent path finding with conflict-based search
Li, J.; Harabor, D.; Stuckey, P. J.; Felner, A.; Ma, H.; and Koenig, S. 2019 · 2019
Cited alongside, same era.
Maven: Multi-agent variational exploration
Mahajan, A.; Rashid, T.; Samvelyan, M.; and Whiteson, S. 2019 · 2019
Cited alongside, same era.
Risk averse robust adversarial reinforcement learning
Pan, X.; Seita, D.; Gao, Y.; and Canny, J. 2019 · 2019
Cited alongside, same era.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Son, K.; Kim, D.; Kang, W. J.; Hostallero, D. E.; and Yi, Y. 2019 · 2019
Cited alongside, same era.
The hanabi challenge: A new frontier for ai research
Bard, N.; Foerster, J. N.; Chandar, S.; Burch, N.; Lanctot, M.; Song, H. F.; Parisotto, E.; Dumoulin, V.; Moitra, S.; Hughes, E.; et al. 2020 · 2020
Cited alongside, same era.
Shared experience actor-critic for multi-agent reinforcement learning
Off-belief learning
Hu, H.; Lerer, A.; Cui, B.; Pineda, L.; Brown, N.; and Foerster, J. 2021 · 2021
Later among the works it cites.
Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Kuba, J. G.; Chen, R.; Wen, M.; Wen, Y.; Sun, F.; Wang, J.; and Yang, Y. 2021 · 2021
Later among the works it cites.
Towards open ad hoc teamwork using graph-based policy learning
Rahman, M. A.; Hopner, N.; Christianos, F.; and Albrecht, S. V. 2021 · 2021
Later among the works it cites.
On the critical role of conventions in adaptive human-AI collaboration
Shih, A.; Sawhney, A.; Kondic, J.; Ermon, S.; and Sadigh, D. 2021 · 2021
Later among the works it cites.
A new formalism, method and open issues for zero-shot coordination
Treutlein, J.; Dennis, M.; Oesterheld, C.; and Foerster, J. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christianos, F.; Schäfer, L.; and Albrecht, S. 2020 · 2020
Cited alongside, same era.
“Other-Play” for Zero-Shot Coordination
Hu, H.; Lerer, A.; Peysakhovich, A.; and Foerster, J. 2020 · 2020
Cited alongside, same era.
K-level Reasoning for Zero-Shot Coordination in Hanabi
Cui, B.; Hu, H.; Pineda, L.; and Foerster, J. 2021 · 2021
Cited alongside, same era.
A deep ensemble method for multi-agent reinforcement learning: A case study on air traffic control
Ghosh, S.; Laguna, S.; Lim, S. H.; Wynter, L.; and Poonawala, H. 2021 · 2021
Cited alongside, same era.
Bøgh, S.; Jensen, P. G.; Kristjansen, M.; Larsen, K. G.; and Nyman, U. 2022 · 2022
Later among the works it cites.
Any-Play: An Intrinsic Augmentation for Zero-Shot Coordination
Lucas, K.; and Allen, R. E. 2022 · 2022
Later among the works it cites.
On-the-fly Strategy Adaptation for ad-hoc Agent Coordination
Zand, J.; Parker-Holder, J.; and Roberts, S. J. 2022 · 2022
Later among the works it cites.
Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward
Sunehag, P.; Lever, G.; Gruslys, A.; Czarnecki, W. M.; Zambaldi, V.; Jaderberg, M.; Lanctot, M.; Sonnerat, N.; Leibo, J. Z.; Tuyls, K.; et al. 2018 · 2087
Closest in time.