Fetching the paper…
Reading the bibliography…
Centralized Training for Decentralized Execution, where agents are trained offline using centralized information but execute in a decentralized manner online, has gained popularity in the multi-agent reinforcement learning community.
A stochastic approximation method
Herbert Robbins and Sutton Monro. 1951 · 1951
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Richard S Sutton. 1985 · 1985
Earlier work this paper cites.
Multi-agent reinforcement learning: independent vs. cooperative agents. In Proceedings of the tenth international conference on machine learning . 330–337
Ming Tan. 1993 · 1993
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier. 1998 · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra. 1998 · 1998
Earlier work this paper cites.
Actor-critic algorithms. In Advances in neural information processing systems . 1008–1014
Vijay R Konda and John N Tsitsiklis. 2000 · 2000
Earlier work this paper cites.
Learning to cooperate via policy search. In Proceedings of the Sixteenth conference on Uncertainty in artificial intelligence . 489–496
Leonid Peshkin, Kee-Eung Kim, Nicolas Meuleau, and Leslie Pack Kaelbling. 2000 · 2000
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games. In UAI . 541–548
Satinder P Singh, Michael J Kearns, and Yishay Mansour. 2000 · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation. In Advances in neural information processing systems . 1057–1063
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour. 2000 · 2000
Earlier work this paper cites.
Convergence of gradient dynamics with a variable learning rate. In International Conference on Machine Learning . 27–34
Michael Bowling and Manuela Veloso. 2001 · 2001
Earlier work this paper cites.
Taming decentralized POMDPs: towards efficient policy computation for multiagent settings. In International Joint Conferences on Artificial Intelligence , Vol. 3. 705–711
Ranjit Nair, Milind Tambe, Makoto Yokoo, David Pynadath, and Stacy Marsella. 2003 · 2003
Earlier work this paper cites.
Predicting and preventing coordination problems in cooperative Q-learning systems.. In International Joint Conferences on Artificial Intelligence , Vol. 2007. 780–785
Nancy Fulda and Dan Ventura. 2007 · 2007
Earlier work this paper cites.
Optimal and approximate Q-value functions for decentralized POMDPs
Frans A Oliehoek, Matthijs TJ Spaan, and Nikos Vlassis. 2008 · 2008
Earlier work this paper cites.
Theoretical advantages of lenient learners: An evolutionary game theoretic perspective
Liviu Panait, Karl Tuyls, and Sean Luke. 2008 · 2008
Earlier work this paper cites.
Incremental policy generation for finite-horizon DEC-POMDPs. In Proceedings of the Nineteenth International Conference on International Conference on Automated Planning and Scheduling . 2–9
Chistopher Amato, Jilles Steeve Dibangoye, and Shlomo Zilberstein. 2009 · 2009
Earlier work this paper cites.
Multi-agent learning with policy prediction. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 24
Chongjie Zhang and Victor Lesser. 2010 · 2010
Earlier work this paper cites.
Independent reinforcement learners in cooperative Markov games: a survey regarding coordination problems
Laetitia Matignon, Guillaume J Laurent, and Nadine Le Fort-Piat. 2012 · 2012
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning. In Advances in neural information processing systems . 2137–2145
Jakob Foerster, Ioannis Alexandros Assael, Nando De Freitas, and Shimon Whiteson. 2016 · 2016
Cited alongside, same era.
Learning to play guess who? and inventing a grounded language as a consequence
Emilio Jorge, Mikael Kågebäck, Fredrik D Johansson, and Emil Gustavsson. 2016 · 2016
Cited alongside, same era.
A Concise Introduction to Decentralized POMDPs
Frans A. Oliehoek and Christopher Amato. 2016 · 2016
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in neural information processing systems . 6379–6390
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Improved cooperative multi-agent reinforcement learning algorithm augmented by mixing demonstrations from centralized policy. In Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems . 1089–1098
Hyun-Rok Lee and Taesik Lee. 2019 · 2019
Later among the works it cites.
Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient
Shihui Li, Yi Wu, Xinyue Cui, Honghua Dong, Fei Fang, and Stuart Russell. 2019 · 2019
Later among the works it cites.
MAVEN: multi-agent variational exploration. In Proceedings of the Thirty-third Annual Conference on Neural Information Processing Systems
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, and Shimon Whiteson. 2019 · 2019
Later among the works it cites.
Learning to teach in cooperative multiagent reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 6128–6136
Shayegan Omidshafiei, Dong-Ki Kim, Miao Liu, Gerald Tesauro, Matthew Riemer, Christopher Amato, Murray Campbell, and Jonathan P How. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel. 2017 · 2017
Cited alongside, same era.
Deep decentralized multi-task multi-Agent reinforcement learning under partial observability. In International Conference on Machine Learning . 2681–2690
Shayegan Omidshafiei, Jason Pazis, Christopher Amato, Jonathan P How, and John Vian. 2017 · 2017
Cited alongside, same era.
Cooperative multi-agent policy gradient. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 459–476
Guillaume Bono, Jilles Steeve Dibangoye, Laëtitia Matignon, Florian Pereyron, and Olivier Simonin. 2018 · 2018
Cited alongside, same era.
Deep multi-Agent reinforcement learning for decentralized continuous cooperative control
Christian Schroeder de Witt, Bei Peng, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Böhmer, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients. In Thirty-second AAAI conference on artificial intelligence
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
Learning attentional communication for multi-agent cooperation. In Advances in neural information processing systems . 7254–7264
Jiechuan Jiang and Zongqing Lu. 2018 · 2018
Cited alongside, same era.
Weighted QMIX: expanding monotonic value function factorisation
Tabish Rashid, Gregory Farquhar, Bei Peng, and Shimon Whiteson. 2018a · 2018
Cited alongside, same era.
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philiph H. S. Torr, Jakob Foerster, and Shimon Whiteson. 2019 · 2019
Later among the works it cites.
QTRAN: learning to factorize with transformation for cooperative multi-agent reinforcement learning. In Proceedings of the 36th International Conference on Machine Learning , Vol. 97. PMLR, 5887–5896
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi. 2019 · 2019
Later among the works it cites.
Achieving cooperation through deep multiagent reinforcement learning in sequential prisoner’s dilemmas. In Proceedings of the First International Conference on Distributed Artificial Intelligence . 1–7
Weixun Wang, Jianye Hao, Yixi Wang, and Matthew Taylor. 2019 · 2019
Later among the works it cites.
Macro-action-based deep multi-agent reinforcement learning. In 3rd Annual Conference on Robot Learning
Yuchen Xiao, Joshua Hoffman, and Christopher Amato. 2019 · 2019
Later among the works it cites.
CM3: cooperative multi-goal multi-stage multi-agent reinforcement learning. In International Conference on Learning Representations
Jiachen Yang, Alireza Nakhaei, David Isele, Kikuo Fujimura, and Hongyuan Zha. 2019 · 2019
Later among the works it cites.
Emergent tool use from multi-agent autocurricula. In Proceedings of the Eighth International Conference on Learning Representations
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch. 2020 · 2020
Later among the works it cites.
Option-critic in cooperative multi-agent systems. In Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems . 1792–1794
Jhelum Chakravorty, Patrick Nadeem Ward, Julien Roy, Maxime Chevalier-Boisvert, Sumana Basu, Andrei Lupu, and Doina Precup. 2020 · 2020
Later among the works it cites.
Likelihood quantile networks for coordinating multi-agent reinforcement learning. In Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems . 798–806
Xueguang Lyu and Christopher Amato. 2020 · 2020
Later among the works it cites.
Multi-agent actor centralized-critic with communication
David Simões, Nuno Lau, and Luís Paulo Reis. 2020 · 2020
Later among the works it cites.
Learning nearly decomposable value functions via communication minimization
Tonghan Wang, Jianhao Wang, Chongyi Zheng, and Chongjie Zhang. 2020b · 2020
Later among the works it cites.
Learning multi-robot decentralized macro-action-based policies via a centralized Q-net. In Proceedings of the International Conference on Robotics and Automation
Yuchen Xiao, Joshua Hoffman, Tian Xia, and Christopher Amato. 2020 · 2020
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning based on team reward. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems . 2085–2087
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2087
Closest in time.