Fetching the paper…
Reading the bibliography…
In many real-world multi-agent cooperative tasks, due to high cost and risk, agents cannot continuously interact with the environment and collect experiences during learning, but have to learn from offline datasets.
Lihong Li, Wei Chu, John Langford, and Robert E Schapire, ‘A Contextual-Bandit Approach To Personalized News Article Recommendation’, in International Conference on World Wide Web (WWW)
2010
Earlier work this paper cites.
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz, ‘Trust Region Policy Optimization’, in International Conference on Machine Learning (ICML)
2015
Earlier work this paper cites.
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra, ‘Continuous Control with Deep Reinforcement Learning.’, in International Conference on Learning Representations (ICLR)
2016
Earlier work this paper cites.
Frans A Oliehoek and Christopher Amato, A Concise Introduction To Decentralized POMDPs
2016
Earlier work this paper cites.
Jakob Foerster, Nantas Nardelli, Gregory Farquhar, Triantafyllos Afouras, Philip HS Torr, Pushmeet Kohli, and Shimon Whiteson, ‘Stabilising Experience Replay for Deep Multi-Agent Reinforcement Learning’, in International Conference on Machine Learning (ICML)
2017
Earlier work this paper cites.
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch, ‘Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments’, in Advances in Neural Information Processing Systems (NeurIPS)
2017
Earlier work this paper cites.
Shayegan Omidshafiei, Jason Pazis, Christopher Amato, Jonathan P How, and John Vian, ‘Deep Decentralized Multi-Task Multi-Agent Reinforcement Learning Under Partial Observability’, in International Conference on Machine Learning (ICML)
2017
Earlier work this paper cites.
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine, ‘Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with A Stochastic Actor’, in International Conference on Machine Learning (ICML)
2018
Earlier work this paper cites.
Gregory Palmer, Karl Tuyls, Daan Bloembergen, and Rahul Savani, ‘Lenient Multi-Agent Deep Reinforcement Learning’, in International Conference on Autonomous Agents and MultiAgent Systems (AAMAS)
2018
Earlier work this paper cites.
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder De Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson, ‘QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning’, in International Conference on Machine Learning (ICML)
2018
Earlier work this paper cites.
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al., ‘Value-Decomposition Networks for Cooperative Multi-Agent Learning Based on Team Reward’, in International Conference on Autonomous Agents and Multiagent Systems (AAMAS)
2018
Earlier work this paper cites.
Scott Fujimoto, David Meger, and Doina Precup, ‘Off-Policy Deep Reinforcement Learning Without Exploration’, in International Conference on Machine Learning (ICML)
2019
Earlier work this paper cites.
Shariq Iqbal and Fei Sha, ‘Actor-Attention-Critic for Multi-Agent Reinforcement Learning’, in International Conference on Machine Learning (ICML)
2019
Earlier work this paper cites.
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine, ‘Stabilizing Off-Policy Q-Learning Via Bootstrapping Error Reduction’, in Advances in Neural Information Processing Systems (NeurIPS)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi, ‘QTRAN: Learning To Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning’, in International Conference on Machine Learning (ICML)
2019
Cited alongside, same era.
Julian Ibarz, Jie Tan, Chelsea Finn, Mrinal Kalakrishnan, Peter Pastor, and Sergey Levine, ‘How To Train Your Robot with Deep Reinforcement Learning: Lessons We Have Learned’, The International Journal of Robotics Research
2021
Closest in time.
2021
Closest in time.
Jianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu, and Chongjie Zhang, ‘Qplex: Duplex Dueling Multi-Agent Q-Learning’, in International Conference on Learning Representations (ICLR)
2021
Closest in time.
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al., ‘Grandmaster Level in StarCraft II Using Multi-Agent Reinforcement Learning’, Nature
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine, ‘Conservative Q-Learning for Offline Reinforcement Learning’, Neural Information Processing Systems (NeurIPS)
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma, ‘MOPO: Model-Based Offline Policy Optimization’, Advances in Neural Information Processing Systems (NeurIPS)
2020
Cited alongside, same era.
Scott Fujimoto and Shixiang Shane Gu, ‘A Minimalist Approach To Offline Reinforcement Learning’, in Thirty-Fifth Conference on Neural Information Processing Systems (NeurIPS)
2021
Cited alongside, same era.
Yiqin Yang, Xiaoteng Ma, Chenghao Li, Zewu Zheng, Qiyuan Zhang, Gao Huang, Jun Yang, and Qianchuan Zhao, ‘Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement Learning’, in Neural Information Processing Systems (NeurIPS)
2021
Closest in time.
2021
Closest in time.
BRAC+: Going Deeper with Behavior Regularized Offline Reinforcement Learning, 2021
Chi Zhang, Sanmukh Rao Kuppannagari, and Viktor Prasanna · 2021
Closest in time.
Tianhao Zhang, Yueheng Li, Chen Wang, Guangming Xie, and Zongqing Lu, ‘FOP: Factorizing Optimal Joint Policy of Maximum-Entropy Multi-Agent Reinforcement Learning’, in International Conference on Machine Learning (ICML)
2021
Closest in time.
Jakub Grudzien Kuba, Ruiqing Chen, Munning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang, ‘Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning’, International Conference on Learning Representations (ICLR)
2022
Closest in time.
Jiafei Lyu, Xiaoteng Ma, Xiu Li, and Zongqing Lu, ‘Mildly Conservative Q-learning for Offline Reinforcement Learning’, Neural Information Processing Systems (NeurIPS)
2022
Closest in time.
Kefan Su and Zongqing Lu, ‘Divergence-Regularized Multi-Agent Actor-Critic’, in International Conference on Machine Learning (ICML)
2022
Closest in time.
Jiangxing Wang, Deheng Ye, and Zongqing Lu, ‘More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy Factorization’, in International Conference on Learning Representations (ICLR)
2023
Closest in time.