Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel model-based multi-agent reinforcement learning approach named Value Decomposition Framework with Disentangled World Model to address the challenge of achieving a common goal of multiple agents interacting in the same environment with reduced sample complexity.
The StarCraft Multi-Agent Challenge, December 2019
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philip H. S. Torr, Jakob Foerster, and Shimon Whiteson · 1902
Earlier work this paper cites.
Multi-Agent Reinforcement Learning: A Selective Overview of Theories and Algorithms, April 2021a
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar · 1911
Earlier work this paper cites.
Dream to Control: Learning Behaviors by Latent Imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 1912
Earlier work this paper cites.
Hierarchical cooperative multi-agent reinforcement learning with skill discovery
Jiachen Yang, Igor Borovikov, and Hongyuan Zha · 1912
Earlier work this paper cites.
Multiagent cooperation and competition with deep reinforcement learning
Ardi Tampuu, Tambet Matiisen, Dorian Kodelja, Ilya Kuzovkin, Kristjan Korjus, Juhan Aru, Jaan Aru, and Raul Vicente · 1932
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
Craig Boutilier · 1996
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent versus cooperative agents
Ming Tan · 1997
Earlier work this paper cites.
A near-optimal poly-time algorithm for learning in a class of stochastic games
Ronen Brafman and Moshe Tennenholtz · 1999
Earlier work this paper cites.
Coordinated reinforcement learning
Carlos Guestrin, Michail G. Lagoudakis, and Ronald E. Parr · 2002
Earlier work this paper cites.
R-max–A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning
Ronen Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
Collaborative multiagent reinforcement learning by payoff propagation
Jelle R. Kok and Nikos Vlassis · 2006
Earlier work this paper cites.
Weighted QMIX: expanding monotonic value function factorisation
Tabish Rashid, Gregory Farquhar, Bei Peng, and Shimon Whiteson · 2006
Earlier work this paper cites.
QPLEX: duplex dueling multi-agent q-learning
Jianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu, and Chongjie Zhang · 2008
Earlier work this paper cites.
Mastering Atari with Discrete World Models, February 2022
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2010
Earlier work this paper cites.
RODE: learning roles to decompose multi-agent tasks
Tonghan Wang, Tarun Gupta, Anuj Mahajan, Bei Peng, Shimon Whiteson, and Chongjie Zhang · 2010
Earlier work this paper cites.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yaodong Yang and Jun Wang · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Manfred Otto Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Earlier work this paper cites.
Continuous deep q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas Kipf and Max Welling · 2016
Cited alongside, same era.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
Landon Kraemer and Bikramjit Banerjee · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
A concise introduction to decentralized POMDPs
Frans A Oliehoek and Christopher Amato · 2016
Cited alongside, same era.
Cooperative multi-agent control using deep reinforcement learning
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Later among the works it cites.
Shared experience actor-critic for multi-agent reinforcement learning
Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht · 2020
Later among the works it cites.
Svqn: Sequential variational soft q-learning networks
Shiyu Huang, Hang Su, Jun Zhu, and Tingling Chen · 2020
Later among the works it cites.
Pma-drl: A parallel model-augmented framework for deep reinforcement learning algorithms
Xufang Luo and Yunhong Wang · 2020
Later among the works it cites.
Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks
Georgios Papoudakis, Filippos Christianos, Lukas Schäfer, and Stefano V Albrecht · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jayesh Gupta, Maxim Egorov, and Mykel J. Kochenderfer · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games
Peng Peng, Quan Yuan, Ying Wen, Yaodong Yang, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
Sebastien Racaniere, Theophane Weber, David P Reichert, Lars Buesing, Arthur Guez, Danilo Rezende, Adria Puigdomenech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Value-Decomposition Networks For Cooperative Multi-Agent Learning, June 2017
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel · 2017
Cited alongside, same era.
Counterfactual Multi-Agent Policy Gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Cited alongside, same era.
Jian Shen, Han Zhao, Weinan Zhang, and Yong Yu · 2020
Later among the works it cites.
Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity
Kaiqing Zhang, Sham Kakade, Tamer Basar, and Lin Yang · 2020
Later among the works it cites.
Learning implicit credit assignment for multi-agent actor-critic
Meng Zhou, Ziyu Liu, Pengwei Sui, Yixuan Li, and Yuk Ying Chung · 2020
Later among the works it cites.
Recurrent independent mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf · 2021
Later among the works it cites.
Efficient model-based multi-agent mean-field reinforcement learning
Barna Pasztor, Ilija Bogunovic, and Andreas Krause · 2021
Later among the works it cites.
Facmac: Factored multi-agent centralised policy gradients
Bei Peng, Tabish Rashid, Christian Schroeder de Witt, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Boehmer, and Shimon Whiteson · 2021
Later among the works it cites.
The surprising effectiveness of MAPPO in cooperative, multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre M. Bayen, and Yi Wu · 2021
Later among the works it cites.
Model-based multi-agent policy optimization with adaptive opponent-wise rollouts
Weinan Zhang, Xihuai Wang, Jian Shen, and Ming Zhou · 2021
Later among the works it cites.
Scalable multi-agent model-based reinforcement learning
Vladimir Egorov and Alexei Shpilman · 2022
Later among the works it cites.
MASER: Multi-agent reinforcement learning with subgoals generated from experience replay buffer
Jeewon Jeon, Woojun Kim, Whiyoung Jung, and Youngchul Sung · 2022
Later among the works it cites.
A Survey on Model-based Reinforcement Learning, June 2022
Fan-Ming Luo, Tian Xu, Hang Lai, Xiong-Hui Chen, Weinan Zhang, and Yang Yu · 2022
Later among the works it cites.
Iso-Dream: Isolating and Leveraging Noncontrollable Visual Dynamics in World Models
Minting Pan, Xiangming Zhu, Yunbo Wang, and Xiaokang Yang · 2022
Later among the works it cites.
Efficient model-based multi-agent reinforcement learning via optimistic equilibrium computation
Pier Giuseppe Sessa, Maryam Kamgarpour, and Andreas Krause · 2022
Later among the works it cites.
Model-based multi-agent reinforcement learning: Recent progress and prospects
Xihuai Wang, Zhicheng Zhang, and Weinan Zhang · 2022
Later among the works it cites.
Blizzard/s2client-proto: Starcraft ii client - protocol definitions used to communicate with starcraft ii
Blizzard · 2023
Closest in time.
Mastering diverse domains through world models
Danijar Hafner, J. Pasukonis, Jimmy Ba, and Timothy P. Lillicrap · 2023
Closest in time.
Boosting multiagent reinforcement learning via permutation invariant and permutation equivariant networks
HAO Jianye, Xiaotian Hao, Hangyu Mao, Weixun Wang, Yaodong Yang, Dong Li, Yan Zheng, and Zhen Wang · 2023
Closest in time.
Contrastive identity-aware learning for multi-agent value decomposition
Shunyu Liu, Yihe Zhou, Jie Song, Tongya Zheng, Kaixuan Chen, Tongtian Zhu, Zunlei Feng, and Mingli Song · 2023
Closest in time.
Models as agents: optimizing multi-step predictions of interactive local models in model-based multi-agent reinforcement learning
Zifan Wu, Chao Yu, Chen Chen, Jianye Hao, and Hankz Hankui Zhuo · 2023
Closest in time.
Self-motivated multi-agent exploration
Shaowei Zhang, Jiahan Cao, Lei Yuan, Yang Yu, and De-Chuan Zhan · 2023
Closest in time.
Imagination-augmented reinforcement learning framework for variable speed limit control
Duo Li and Joan Lasenby · 2024
Closest in time.