Fetching the paper…
Reading the bibliography…
In many coordination problems, independently reasoning humans are able to discover mutually compatible policies.
Theory of Games and Economic Behavior
von Neumann, J. and Morgenstern, O · 1947
Earlier work this paper cites.
Non-cooperative games
Nash, J. F · 1951
Earlier work this paper cites.
A further generalization of the Kakutani fixed point theorem, with application to Nash equilibrium points
Glicksberg, I. L · 1952
Earlier work this paper cites.
Extensive games and the problem of information
Kuhn, H. W · 1953
Earlier work this paper cites.
The tracing procedure: a Bayesian approach to defining a solution for n-person noncooperative games
Harsanyi, J. C · 1975
Earlier work this paper cites.
The Strategy of Conflict
Schelling, T. C · 1980
Earlier work this paper cites.
A general theory of equilibrium selection in games
Harsanyi, J. C. and Selten, R · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1988
Earlier work this paper cites.
Probability with Martingales
Williams, D · 1991
Earlier work this paper cites.
Game theory for applied economists
Gibbons, R. S · 1992
Earlier work this paper cites.
The nature of salience: An experimental investigation of pure coordination games
Mehta, J., Starmer, C., and Sugden, R · 1994
Earlier work this paper cites.
A course in game theory
Osborne, M. J. and Rubinstein, A · 1994
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G · 1994
Earlier work this paper cites.
Sequential optimality and coordination in multiagent systems
Boutilier, C · 1999
Earlier work this paper cites.
The canonical extensive form of a game form: Symmetries
Peleg, B., Rosenmüller, J., and Sudhölter, P · 1999
Earlier work this paper cites.
Weak isomorphisms of extensive games
Casajus, A · 2001
Earlier work this paper cites.
Deep Blue
Campbell, M., Hoane Jr, A. J., and Hsu, F.-h · 2002
Earlier work this paper cites.
Computation of the Nash equilibrium selected by the tracing procedure in n-person games
Herings, P. J.-J. and Van Den Elzen, A · 2002
Cited alongside, same era.
Equilibrium selection in stochastic games
Herings, P. J.-J. and Peeters, R. J · 2003
Cited alongside, same era.
Taming decentralized POMDPs: Towards efficient policy computation for multiagent settings
Nair, R., Tambe, M., Yokoo, M., Pynadath, D., and Marsella, S · 2003
Cited alongside, same era.
Dec-POMDPs and extensive form games: equivalence of models and algorithms
Oliehoek, F., Vlassis, N., et al · 2006
Cited alongside, same era.
Learning to communicate in a decentralized environment
Goldman, C. V., Allen, M., and Zilberstein, S · 2007
Cited alongside, same era.
Convention: A philosophical study
Lewis, D · 2008
Cited alongside, same era.
A concise introduction to decentralized POMDPs
Oliehoek, F. A., Amato, C., et al · 2016
Later among the works it cites.
Policy gradient with value function approximation for collective multiagent planning
Nguyen, D. T., Kumar, A., and Lau, H. C · 2017
Later among the works it cites.
Equivalence between policy gradients and soft Q-learning
Schulman, J., Abbeel, P., and Chen, X · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D · 2017
Later among the works it cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimal and approximate Q-value functions for decentralized POMDPs
Oliehoek, F. A., Spaan, M. T., and Vlassis, N · 2008
Cited alongside, same era.
Causality: Models, Reasoning, and Inference
Pearl, J · 2009
Cited alongside, same era.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
Stone, P., Kaminka, G., Kraus, S., and Rosenschein, J · 2010
Cited alongside, same era.
Empirical evaluation of ad hoc teamwork in the pursuit domain
Barrett, S., Stone, P., and Kraus, S · 2011
Cited alongside, same era.
Exploiting symmetries for single-and multi-agent partially observable stochastic domains
Kang, B. K. and Kim, K.-E · 2012
Cited alongside, same era.
An Introduction to the Theory of Groups
Rotman, J. J · 2012
Cited alongside, same era.
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., and Graepel, T · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
On the utility of learning about humans for human-AI coordination
Carroll, M., Shah, R., Ho, M. K., Griffiths, T., Seshia, S., Abbeel, P., and Dragan, A · 2019
Later among the works it cites.
Bayesian action decoder for deep multi-agent reinforcement learning
Foerster, J., Song, F., Hughes, E., Burch, N., Dunning, I., Whiteson, S., Botvinick, M., and Bowling, M · 2019
Later among the works it cites.
Learning existing social conventions via observationally augmented self-play
Lerer, A. and Peysakhovich, A · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
The StarCraft Multi-Agent Challenge
Samvelyan, M., Rashid, T., de Witt, C. S., Farquhar, G., Nardelli, N., Rudner, T. G. J., Hung, C.-M., Torr, P. H. S., Foerster, J., and Whiteson, S · 2019
Later among the works it cites.
“Other-play” for zero-shot coordination
Hu, H., Lerer, A., Peysakhovich, A., and Foerster, J · 2020
Later among the works it cites.
Adversarially guided self-play for adopting social conventions
Tucker, M., Zhou, Y., and Shah, J · 2020
Later among the works it cites.
Symmetry, equilibria, and robustness in common-payoff games
Emmons, S., Oesterheld, C., Critch, A., Conitzer, V., and Russell, S · 2021
Closest in time.