Fetching the paper…
Reading the bibliography…
In cooperative multi-agent reinforcement learning, a collection of agents learns to interact in a shared environment to achieve a common goal.
Multi-Agent Reinforcement Learning: A Selective Overview of Theories and Algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar. 2019 · 1911
Earlier work this paper cites.
A linear algorithm for testing equivalence of finite automata . Vol. 114
John E Hopcroft. 1971 · 1971
Earlier work this paper cites.
On observability of discrete-event systems
Feng Lin and Walter Murray Wonham. 1988 · 1988
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan. 1992 · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents. In Proceedings of the 10th International Conference on Machine Learning . 330–337
Ming Tan. 1993 · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman. 1994 · 1994
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes. In Proceedings of the 6th conference on Theoretical aspects of rationality and knowledge . Morgan Kaufmann Publishers Inc., 195–210
Craig Boutilier. 1996 · 1996
Earlier work this paper cites.
On the complexity of projections of discrete-event systems. In Proceedings of the International Workshop of Discrete Event Systems . Citeseer, 201–206
K Wong. 1998 · 1998
Earlier work this paper cites.
Hierarchical multi-agent reinforcement learning. In Proceedings of the 5th International Conference on Autonomous Agents . 246–253
Rajbala Makar, Sridhar Mahadevan, and Mohammad Ghavamzadeh. 2001 · 2001
Earlier work this paper cites.
Coordinated reinforcement learning. In International Conference on Machine Learning , Vol. 2. Citeseer, 227–234
Carlos Guestrin, Michail Lagoudakis, and Ronald Parr. 2002 · 2002
Earlier work this paper cites.
Extended Markov Games to Learn Multiple Tasks in Multi-Agent Reinforcement Learning
Borja G. León and Francesco Belardinelli. 2020 · 2002
Earlier work this paper cites.
Using the max-plus algorithm for multiagent decision making in coordination graphs. In Robot Soccer World Cup . Springer, 1–12
Jelle R Kok and Nikos Vlassis. 2005 · 2005
Earlier work this paper cites.
Hierarchical multi-agent reinforcement learning
Mohammad Ghavamzadeh, Sridhar Mahadevan, and Rajbala Makar. 2006 · 2006
Earlier work this paper cites.
Principles of model checking
Christel Baier and Joost-Pieter Katoen. 2008 · 2008
Earlier work this paper cites.
Introduction to discrete event systems
Christos G Cassandras and Stephane Lafortune. 2009 · 2009
Cited alongside, same era.
Guaranteed global performance through local coordinations
Mohammad Karimadini and Hai Lin. 2011 · 2011
Cited alongside, same era.
Modeling and control of logical discrete event systems . Vol. 300
Ratnesh Kumar and Vijay K Garg. 2012 · 2012
Cited alongside, same era.
Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
Laetitia Matignon, Guillaume J Laurent, and Nadine Le Fort-Piat. 2012 · 2012
Cited alongside, same era.
Checking NFA equivalence with bisimulations up to congruence
Filippo Bonchi and Damien Pous. 2013 · 2013
Cited alongside, same era.
Automatic synthesis of cooperative multi-agent systems. In 53rd IEEE Conference on Decision and Control . IEEE, 6173–6178
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Later among the works it cites.
LTL and beyond: Formal languages for reward function specification in reinforcement learning. In Proceedings of the 28th International Joint Conference on Artificial Intelligence . 6065–6073
Alberto Camacho, R Toro Icarte, Toryn Q Klassen, Richard Valenzano, and Sheila A McIlraith. 2019 · 2019
Later among the works it cites.
A survey and critique of multiagent deep reinforcement learning
Pablo Hernandez-Leal, Bilal Kartal, and Matthew E Taylor. 2019b · 2019
Later among the works it cites.
Learning Reward Machines for Partially Observable Reinforcement Learning. In Advances in Neural Information Processing Systems . 15497–15508
Rodrigo Toro Icarte, Ethan Waldie, Toryn Klassen, Rick Valenzano, Margarita Castro, and Sheila McIlraith. 2019 · 2019
Later among the works it cites.
Maven: Multi-agent variational exploration. In Advances in Neural Information Processing Systems . 7613–7624
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, and Shimon Whiteson. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jin Dai and Hai Lin. 2014 · 2014
Cited alongside, same era.
Decentralized hybrid formation control of unmanned aerial vehicles. In American Control Conference . IEEE, 3887–3892
Ali Karimoddini, Mohammad Karimadini, and Hai Lin. 2014 · 2014
Cited alongside, same era.
Comparing exploration strategies for q-learning in random stochastic mazes. In 2016 IEEE Symposium Series on Computational Intelligence (SSCI) . IEEE, 1–8
Arryon D Tijsma, Madalina M Drugan, and Marco A Wiering. 2016 · 2016
Cited alongside, same era.
Coordinated deep reinforcement learners for traffic light control
Elise Van der Pol and Frans A Oliehoek. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning. In Proceedings of the 13th AAAI Conference on Artificial Intelligence
Hado Van Hasselt, Arthur Guez, and David Silver. 2016 · 2016
Cited alongside, same era.
Using reward machines for high-level task specification and decomposition in reinforcement learning. In International Conference on Machine Learning . 2112–2121
Rodrigo Toro Icarte, Toryn Klassen, Richard Valenzano, and Sheila McIlraith. 2018 · 2018
Cited alongside, same era.
Enforcing signal temporal logic specifications in multi-agent adversarial environments: A deep Q-learning approach. In 2018 IEEE Conference on Decision and Control (CDC) . IEEE, 4141–4146
Devaprakash Muniraj, Kyriakos G Vamvoudakis, and Mazen Farhood. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
The StarCraft Multi-Agent Challenge. In Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems . 2186–2188
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim GJ Rudner, Chia-Man Hung, Philip HS Torr, Jakob Foerster, and Shimon Whiteson. 2019 · 2019
Later among the works it cites.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning. In International Conference on Machine Learning . PMLR, 5887–5896
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi. 2019 · 2019
Later among the works it cites.
Hierarchical Deep Multiagent Reinforcement Learning with Temporal Abstraction
Hongyao Tang, Jianye Hao, Tangjie Lv, Yingfeng Chen, Zongzhang Zhang, Hangtian Jia, Chunxu Ren, Yan Zheng, Zhaopeng Meng, Changjie Fan, and Li Wang. 2019 · 2019
Later among the works it cites.
Policy Synthesis for Factored MDPs with Graph Temporal Logic Specifications. In Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems . 267–275
Murat Cubuktepe, Zhe Xu, and Ufuk Topcu. 2020 · 2020
Closest in time.
Probabilistic Swarm Guidance with Graph Temporal Logic Specifications. In Proceedings of Robotics: Science and Systems XVI
Franck Djeumou, Zhe Xu, and Ufuk Topcu. 2020 · 2020
Closest in time.
Joint inference of reward machines and policies for reinforcement learning. In Proceedings of the International Conference on Automated Planning and Scheduling , Vol. 30. 590–598
Zhe Xu, Ivan Gavran, Yousef Ahmad, Rupak Majumdar, Daniel Neider, Ufuk Topcu, and Bo Wu. 2020 · 2020
Closest in time.
Cooperative tasking for deterministic specification automata
Mohammad Karimadini, Hai Lin, and Ali Karimoddini. 2016 · 2087
Closest in time.
Value-decomposition networks for cooperative multi-agent learning based on team reward. In Proceedings of the 17th International Conference on Autonomous Agents and Multiagent Systems . 2085–2087
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2087
Closest in time.