Fetching the paper…
Reading the bibliography…
In order for artificial agents to coordinate effectively with people, they must act consistently with existing conventions (e.g.
Zur theorie der gesellschaftsspiele
J von Neumann. 1928 · 1928
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley. 1953 · 1953
Earlier work this paper cites.
Convention: A philosophical study
David Lewis. 1969 · 1969
Earlier work this paper cites.
Multimarket oligopoly: Strategic substitutes and complements
Jeremy I Bulow, John D Geanakoplos, and Paul D Klemperer. 1985 · 1985
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman. 1994 · 1994
Earlier work this paper cites.
Understanding the Emergence of Conventions in Multi-Agent Systems.. In ICMAS , Vol. 95. 384–389
Adam Walker and Michael Wooldridge. 1995 · 1995
Earlier work this paper cites.
On the emergence of social conventions: modeling, analysis, and simulations
Yoav Shoham and Moshe Tennenholtz. 1997 · 1997
Earlier work this paper cites.
The theory of learning in games . Vol. 2
Drew Fudenberg and David K Levine. 1998 · 1998
Earlier work this paper cites.
Scaling reinforcement learning toward RoboCup soccer. In ICML , Vol. 1. Citeseer, 537–544
Peter Stone and Richard S Sutton. 2001 · 2001
Earlier work this paper cites.
Reinforcement learning of coordination in cooperative multi-agent systems
Spiros Kapetanakis and Daniel Kudenko. 2002b · 2002
Earlier work this paper cites.
Potential-based shaping and Q-value initialization are equivalent
Eric Wiewiora. 2003 · 2003
Earlier work this paper cites.
The grammar of society: The nature and dynamics of social norms
Cristina Bicchieri. 2005 · 2005
Earlier work this paper cites.
Social reward shaping in the prisoner’s dilemma. In Proceedings of the 7th international joint conference on Autonomous agents and multiagent systems-Volume 3 . International Foundation for Autonomous Agents and Multiagent Systems, 1389–1392
Monica Babes, Enrique Munoz De Cote, and Michael L Littman. 2008 · 2008
Earlier work this paper cites.
Game theory of mind
Wako Yoshida, Ray J Dolan, and Karl J Friston. 2008 · 2008
Cited alongside, same era.
Ad hoc autonomous agent teams: Collaboration without pre-coordination. In Twenty-Fourth AAAI Conference on Artificial Intelligence
Peter Stone, Gal A Kaminka, Sarit Kraus, and Jeffrey S Rosenschein. 2010 · 2010
Cited alongside, same era.
Theoretical considerations of potential-based reward shaping for multi-agent systems. In The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 1 . International Foundation for Autonomous Agents and Multiagent Systems, 225–232
Sam Devlin and Daniel Kudenko. 2011 · 2011
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
A unified game-theoretic approach to multiagent reinforcement learning. In Advances in Neural Information Processing Systems . 4193–4206
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Julien Perolat, David Silver, Thore Graepel, et al · 2017
Later among the works it cites.
Multi-agent cooperation and the emergence of (natural) language. In International Conference on Learning Representations
Angeliki Lazaridou, Alexander Peysakhovich, and Marco Baroni. 2017 · 2017
Later among the works it cites.
Multi-agent Reinforcement Learning in Sequential Social Dilemmas. In Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems . International Foundation for Autonomous Agents and Multiagent Systems, 464–473
Joel Z Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel. 2017 · 2017
Later among the works it cites.
Maintaining cooperation in complex social dilemmas using deep reinforcement learning
Adam Lerer and Alexander Peysakhovich. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning. In Advances in Neural Information Processing Systems . 2137–2145
Jakob Foerster, Yannis Assael, Nando de Freitas, and Shimon Whiteson. 2016 · 2016
Cited alongside, same era.
Learning to Play Guess Who? and Inventing a Grounded Language as a Consequence
Emilio Jorge, Mikael Kågebäck, and Emil Gustavsson. 2016 · 2016
Cited alongside, same era.
Coordinate to cooperate or compete: abstract goals and joint intentions in social interaction. In Proceedings of the 38th annual conference of the cognitive science society
Max Kleiman-Weiner, MK Ho, JL Austerweil, Michael L Littman, and Josh B Tenenbaum. 2016 · 2016
Cited alongside, same era.
Learning multiagent communication with backpropagation. In Advances in Neural Information Processing Systems . 2244–2252
Sainbayar Sukhbaatar, Rob Fergus, et al · 2016
Cited alongside, same era.
Making friends on the fly: Cooperating with new teammates
Samuel Barrett, Avi Rosenfeld, Sarit Kraus, and Peter Stone. 2017 · 2017
Cited alongside, same era.
Learning with Opponent-Learning Awareness
Jakob N Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Deep Q-learning from Demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Gabriel Dulac-Arnold, et al · 2017
Cited alongside, same era.
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Later among the works it cites.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel. 2017 · 2017
Later among the works it cites.
Consequentialist conditional cooperation in social dilemmas with imperfect information
Alexander Peysakhovich and Adam Lerer. 2017a · 2017
Later among the works it cites.
Prosocial learning agents solve generalized Stag Hunts better than selfish ones
Alexander Peysakhovich and Adam Lerer. 2017b · 2017
Later among the works it cites.
Locally noisy autonomous agents improve global human coordination in network experiments
Hirokazu Shirado and Nicholas A Christakis. 2017 · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
PyTorch Implementations of Asynchronous Advantage Actor Critic
Ilya Kostrikov. 2018 · 2018
Closest in time.
Cinjon Resnick, Ilya Kulikov, Kyunghyun Cho, and Jason Weston. 2018 · 2018
Closest in time.