Fetching the paper…
Reading the bibliography…
Cooperative Multi-Agent Reinforcement Learning (MARL) algorithms, trained only to optimize task reward, can lead to a concentration of power where the failure or adversarial intent of a single agent could decimate the reward of every agent in the system.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman. 1994 · 1994
Earlier work this paper cites.
Friend-or-foe Q-learning in general-sum games. In ICML , Vol. 1. 322–328
Michael L Littman et al · 2001
Earlier work this paper cites.
Responsibility and blame: A structural-model approach
Hana Chockler and Joseph Y Halpern. 2004 · 2004
Earlier work this paper cites.
Coordinating tasks in agent organizations. In International Workshop on Coordination, Organizations, Institutions, and Norms in Agent Systems . Springer, 32–47
Virginia Dignum and Frank Dignum. 2006 · 2006
Earlier work this paper cites.
Structural aspects of the evaluation of agent organizations. In International Workshop on Coordination, Organizations, Institutions, and Norms in Agent Systems . Springer, 3–18
Davide Grossi, Frank Dignum, Virginia Dignum, Mehdi Dastani, and Làmber Royakkers. 2006 · 2006
Earlier work this paper cites.
Ad Hoc Autonomous Agent Teams: Collaboration without Pre-Coordination. In Proceedings of the Twenty-Fourth Conference on Artificial Intelligence
Peter Stone, Gal A. Kaminka, Sarit Kraus, and Jeffrey S. Rosenschein. 2010 · 2010
Earlier work this paper cites.
Empirical evaluation of ad hoc teamwork in the pursuit domain. In The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 2 . 567–574
Samuel Barrett, Peter Stone, and Sarit Kraus. 2011 · 2011
Earlier work this paper cites.
The sources of social power: volume 1, a history of power from the beginning to AD 1760 . Vol. 1
Michael Mann. 2012 · 2012
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
Responsibility judgments in voting scenarios.. In CogSci
Tobias Gerstenberg, Joseph Y Halpern, and Joshua B Tenenbaum. 2015 · 2015
Earlier work this paper cites.
Learning with opponent-learning awareness
Jakob N Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Adversarial attacks on neural network policies
Sandy Huang, Nicolas Papernot, Ian Goodfellow, Yan Duan, and Pieter Abbeel. 2017 · 2017
Cited alongside, same era.
Delving into adversarial attacks on deep policies
Jernej Kos and Dawn Song. 2017 · 2017
Cited alongside, same era.
Multi-agent reinforcement learning in sequential social dilemmas
Joel Z Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel. 2017 · 2017
Cited alongside, same era.
Tactics of adversarial attack on deep reinforcement learning agents
Adversarial policies: Attacking deep reinforcement learning
Adam Gleave, Michael Dennis, Cody Wild, Neel Kant, Sergey Levine, and Stuart Russell. 2019 · 2019
Later among the works it cites.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning. In International conference on machine learning . PMLR, 3040–3049
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro Ortega, DJ Strouse, Joel Z Leibo, and Nando De Freitas. 2019 · 2019
Later among the works it cites.
Finding friend and foe in multi-agent games
Jack Serrino, Max Kleiman-Weiner, David C Parkes, and Josh Tenenbaum. 2019 · 2019
Later among the works it cites.
Optimal Policies Tend to Seek Power
Alexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch, and Prasad Tadepalli. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yen-Chen Lin, Zhang-Wei Hong, Yuan-Hong Liao, Meng-Li Shih, Ming-Yu Liu, and Min Sun. 2017 · 2017
Cited alongside, same era.
Towards formal definitions of blameworthiness, intention, and moral responsibility. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32
Joseph Halpern and Max Kleiman-Weiner. 2018 · 2018
Cited alongside, same era.
Inequity aversion improves cooperation in intertemporal social dilemmas
Edward Hughes, Joel Z Leibo, Matthew Phillips, Karl Tuyls, Edgar Dueñez-Guzman, Antonio García Castañeda, Iain Dunning, Tina Zhu, Kevin McKee, Raphael Koster, et al · 2018
Cited alongside, same era.
On the utility of learning about humans for human-ai coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan. 2019 · 2019
Cited alongside, same era.
Blameworthiness in multi-agent settings. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 525–532
Meir Friedenberg and Joseph Y Halpern. 2019 · 2019
Cited alongside, same era.
Conservative agency via attainable utility preservation. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society . 385–391
Alexander Matt Turner, Dylan Hadfield-Menell, and Prasad Tadepalli. 2020a
Cited in the paper.
Natasha Alechina, Joseph Y Halpern, and Brian Logan. 2020 · 2020
Later among the works it cites.
“Other-Play” for Zero-Shot Coordination. In International Conference on Machine Learning . PMLR, 4399–4410
Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob Foerster. 2020 · 2020
Later among the works it cites.
Avoiding side effects in complex environments
Alex Turner, Neale Ratzlaff, and Prasad Tadepalli. 2020b · 2020
Later among the works it cites.
Power: A radical view
Steven Lukes. 2021 · 2021
Later among the works it cites.
A new formalism, method and open issues for zero-shot coordination. In International Conference on Machine Learning . PMLR, 10413–10423
Johannes Treutlein, Michael Dennis, Caspar Oesterheld, and Jakob Foerster. 2021 · 2021
Later among the works it cites.