Fetching the paper…
Reading the bibliography…
Training multiple agents to coordinate is an essential problem with applications in robotics, game theory, economics, and social sciences.
Distributed intelligence for air fleet control
Randall Steeb, Stephanie Cammarata, Frederick A Hayes-Roth, Perry W Thorndyke, and Robert E Wesson. 1981 · 1981
Earlier work this paper cites.
Strategic information transmission
Vincent P Crawford and Joel Sobel. 1982 · 1982
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton. 1990 · 1990
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Dean A Pomerleau. 1991 · 1991
Earlier work this paper cites.
Packet routing in dynamically changing networks: A reinforcement learning approach
Justin Boyan and Michael Littman. 1993 · 1993
Earlier work this paper cites.
Sophisticated and distributed: The transportation domain. In Proceedings of 9th IEEE Conference on Artificial Intelligence for Applications . IEEE, 454
K Fischer, N Kuhn, HJ Muller, JP Muller, and M Pischel. 1993 · 1993
Earlier work this paper cites.
An agent architecture for distributed medical care. In Intelligent Agents: ECAI-94 Workshop on Agent Theories, Architectures, and Languages Amsterdam, The Netherlands August 8–9, 1994 Proceedings 1 . Springer, 219–232
Jun Huang, Nicholas R Jennings, and John Fox. 1995 · 1994
Earlier work this paper cites.
Integrating intelligent systems into a cooperating community for electricity distribution management
LászlóZ Varga, Nick R Jennings, and David Cockburn. 1994 · 1994
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes. In TARK , Vol. 96. Citeseer, 195–210
Craig Boutilier. 1996 · 1996
Earlier work this paper cites.
Emergent actors in world politics: how states and nations develop and dissolve . Vol. 2
Lars-Erik Cederman. 1997 · 1997
Earlier work this paper cites.
Creatures: Artificial life autonomous software agents for home entertainment. In Proceedings of the first international conference on Autonomous agents . 22–29
Stephen Grand, Dave Cliff, and Anil Malhotra. 1997 · 1997
Earlier work this paper cites.
Multi-machine scheduling-a multi-agent learning approach. In Proceedings International Conference on Multi Agent Systems (Cat. No. 98EX160) . IEEE, 42–48
Wilfried Brauer and Gerhard Weiß. 1998 · 1998
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier. 1998 · 1998
Earlier work this paper cites.
Friend-or-foe Q-learning in general-sum games. In ICML , Vol. 1. 322–328
Michael L Littman et al · 2001
Earlier work this paper cites.
Efficient learning equilibrium
Ronen Brafman and Moshe Tennenholtz. 2002 · 2002
Earlier work this paper cites.
Coordination in multiagent reinforcement learning: A Bayesian approach. In Proceedings of the second international joint conference on Autonomous agents and multiagent systems . 709–716
Georgios Chalkiadakis and Craig Boutilier. 2003 · 2003
Earlier work this paper cites.
Taming decentralized POMDPs: Towards efficient policy computation for multiagent settings. In IJCAI , Vol. 3. Citeseer, 705–711
Ranjit Nair, Milind Tambe, Makoto Yokoo, David Pynadath, and Stacy Marsella. 2003 · 2003
Earlier work this paper cites.
Multiagent traffic management: A reservation-based intersection control mechanism. In Autonomous Agents and Multiagent Systems, International Joint Conference on , Vol. 3. Citeseer, 530–537
Kurt Dresner and Peter Stone. 2004 · 2004
Earlier work this paper cites.
Cooperative multi-agent learning: The state of the art
Liviu Panait and Sean Luke. 2005 · 2005
Earlier work this paper cites.
Ad hoc autonomous agent teams: Collaboration without pre-coordination. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 24. 1504–1509
Peter Stone, Gal Kaminka, Sarit Kraus, and Jeffrey Rosenschein. 2010 · 2010
Earlier work this paper cites.
Empirical evaluation of ad hoc teamwork in the pursuit domain. In The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 2 . 567–574
Samuel Barrett, Peter Stone, and Sarit Kraus. 2011 · 2011
Earlier work this paper cites.
Coordinating multi-agent reinforcement learning with limited communication. In Proceedings of the 2013 international conference on Autonomous agents and multi-agent systems . 1101–1108
Chongjie Zhang and Victor Lesser. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations, ICLR 2015
Dzmitry Bahdanau, Kyung Hyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. 2015 · 2015
Cited alongside, same era.
Multi-Agent Cooperation and the Emergence of (Natural) Language. In International Conference on Learning Representations
Angeliki Lazaridou, Alexander Peysakhovich, and Marco Baroni. 2017 · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems . 6379–6390
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Adversarially guided self-play for adopting social conventions
Mycal Tucker, Yilun Zhou, and Julie Shah. 2020 · 2020
Later among the works it cites.
Learning to interactively learn and assist. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 2535–2543
Mark Woodward, Chelsea Finn, and Karol Hausman. 2020 · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma. 2020 · 2020
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu. 2021 · 2021
Later among the works it cites.
Offline decentralized multi-agent reinforcement learning
Jiechuan Jiang and Zongqing Lu. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Counterfactual multi-agent policy gradients. In Proceedings of the AAAI conference on artificial intelligence , Vol. 32
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
David Ha and Jürgen Schmidhuber. 2018 · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning . PMLR, 1861–1870
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning. In International Conference on Machine Learning . PMLR, 4295–4304
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
On the utility of learning about humans for human-ai coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan. 2019 · 2019
Cited alongside, same era.
Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning. In International Conference on Machine Learning . 3040–3049
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro Ortega, Dj Strouse, Joel Z Leibo, and Nando De Freitas. 2019 · 2019
Cited alongside, same era.
Learning existing social conventions via observationally augmented self-play. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society . 107–114
Adam Lerer and Alexander Peysakhovich. 2019 · 2019
Cited alongside, same era.
Offline Reinforcement Learning with Implicit Q-Learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine. 2021 · 2021
Later among the works it cites.
Online multi-agent reinforcement learning for decentralized inverter-based volt-var control
Haotian Liu and Wenchuan Wu. 2021 · 2021
Later among the works it cites.
Contrasting centralized and decentralized critics in multi-agent reinforcement learning
Xueguang Lyu, Yuchen Xiao, Brett Daley, and Christopher Amato. 2021 · 2021
Later among the works it cites.
Facmac: Factored multi-agent centralised policy gradients
Bei Peng, Tabish Rashid, Christian Schroeder de Witt, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Böhmer, and Shimon Whiteson. 2021 · 2021
Later among the works it cites.
Offline Reinforcement Learning for Autonomous Driving with Safety and Exploration Enhancement
Tianyu Shi, Dong Chen, Kaian Chen, and Zhaojian Li. 2021 · 2021
Later among the works it cites.
Feedback in imitation learning: The three regimes of covariate shift
Jonathan Spencer, Sanjiban Choudhury, Arun Venkatraman, Brian Ziebart, and J Andrew Bagnell. 2021 · 2021
Later among the works it cites.
Offline reinforcement learning with reverse model-based imagination
Jianhao Wang, Wenzhe Li, Haozhe Jiang, Guangxiang Zhu, Siyuan Li, and Chongjie Zhang. 2021 · 2021
Later among the works it cites.
MAMBPO: Sample-efficient multi-robot reinforcement learning using learned world models. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 5635–5640
Daniël Willemsen, Mario Coppola, and Guido CHE de Croon. 2021 · 2021
Later among the works it cites.
Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement Learning
Yiqin Yang, Xiaoteng Ma, Chenghao Li, Zewu Zheng, Qiyuan Zhang, Gao Huang, Jun Yang, and Qianchuan Zhao. 2021 · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn. 2021 · 2021
Later among the works it cites.
Model-based Multi-agent Policy Optimization with Adaptive Opponent-wise Rollouts
Weinan Zhang, Xihuai Wang, Jian Shen, and Ming Zhou. 2021 · 2021
Later among the works it cites.
Learning to Guide and to be Guided in the Architect-Builder Problem. In International Conference on Learning Representations
Paul Barde, Tristan Karch, Derek Nowrouzezahrai, Clément Moulin-Frier, Christopher Pal, and Pierre-Yves Oudeyer. 2022 · 2022
Later among the works it cites.
Plan better amid conservatism: Offline multi-agent reinforcement learning with actor rectification. In International Conference on Machine Learning . PMLR, 17221–17237
Ling Pan, Longbo Huang, Tengyu Ma, and Huazhe Xu. 2022 · 2022
Later among the works it cites.
Offline Multi-Agent Reinforcement Learning with Knowledge Distillation
Wei-Cheng Tseng, Tsun-Hsuan Johnson Wang, Yen-Chen Lin, and Phillip Isola. 2022 · 2022
Later among the works it cites.
Eugene Vinitsky, Nathan Lichtlé, Xiaomeng Yang, Brandon Amos, and Jakob Foerster. 2022 · 2022
Later among the works it cites.
The surprising effectiveness of ppo in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. 2022 · 2022
Later among the works it cites.
Centralized Model and Exploration Policy for Multi-Agent RL. In Proceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems (Virtual Event, New Zealand) (AAMAS ’22) . International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 1500–1508
Qizhen Zhang, Chris Lu, Animesh Garg, and Jakob Foerster. 2022 · 2022
Later among the works it cites.
Off-policy deep reinforcement learning without exploration. In International conference on machine learning . PMLR, 2052–2062
Scott Fujimoto, David Meger, and Doina Precup. 2019 · 2062
Closest in time.