Fetching the paper…
Reading the bibliography…
Training a multi-agent reinforcement learning (MARL) algorithm is more challenging than training a single-agent reinforcement learning algorithm, because the result of a multi-agent task strongly depends on the complex interactions among agents and their interactions with a stochastic and dynamic environment.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman. 1994 · 1994
Earlier work this paper cites.
Multiagent reinforcement learning: theoretical framework and an algorithm.. In Proceedings of the Fifteenth International Conference on Machine Learning , Vol. 98. Citeseer, 242–250
Junling Hu, Michael P Wellman, et al · 1998
Earlier work this paper cites.
On amount and quality of bias in reinforcement learning. In IEEE SMC’99 Conference Proceedings. 1999 IEEE International Conference on Systems, Man, and Cybernetics , Vol. 2. IEEE, 728–733
G Hailu and G Sommer. 1999 · 1999
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems. In In Proceedings of the Seventeenth International Conference on Machine Learning . Citeseer
Martin Lauer and Martin Riedmiller. 2000 · 2000
Earlier work this paper cites.
Friend-or-Foe Q-Learning in General-Sum Games. In Proceedings of the Eighteenth International Conference on Machine Learning , Vol. 1. Morgan Kaufmann Publishers Inc., 322–328
Michael L. Littman. 2001 · 2001
Earlier work this paper cites.
Deterministic Policy Gradient Algorithms. In International Conference on Machine Learning , Vol. 32. PMLR, 387–395
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016 · 2016
Cited alongside, same era.
Hindsight Experience Replay. In Advances in Neural Information Processing Systems . 5048–5058
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba. 2017 · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems . 6379–6390
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Cited alongside, same era.
Exploration with unreliable intrinsic reward in multi-agent reinforcement learning
Wendelin Böhmer, Tabish Rashid, and Shimon Whiteson. 2019 · 2019
Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 4213–4220
Shihui Li, Yi Wu, Xinyue Cui, Honghua Dong, Fei Fang, and Stuart Russell. 2019 · 2019
Later among the works it cites.
Hindsight policy gradients
Paulo Rauber, Avinash Ummadisingu, Filipe Mutz, and Juergen Schmidhuber. 2019 · 2019
Later among the works it cites.
Promoting Coordination through Policy Regularization in Multi-Agent Deep Reinforcement Learning
Julien Roy, Paul Barde, Félix G Harvey, Derek Nowrouzezahrai, and Christopher Pal. 2019 · 2019
Later among the works it cites.
Multi-Agent Actor-Critic with Hierarchical Graph Attention Network. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 7236–7243
Heechang Ryu, Hayong Shin, and Jinkyoo Park. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Coordinated Exploration via Intrinsic Rewards for Multi-Agent Reinforcement Learning
Shariq Iqbal and Fei Sha. 2019b · 2019
Cited alongside, same era.
Actor-attention-critic for multi-agent reinforcement learning. In International Conference on Machine Learning , Vol. 97. PMLR, 2961–2970
Shariq Iqbal and Fei Sha. 2019a
Cited in the paper.
Tonghan Wang, Jianhao Wang, Yi Wu, and Chongjie Zhang. 2020 · 2020
Later among the works it cites.