Fetching the paper…
Reading the bibliography…
Humans are capable of attributing latent mental contents such as beliefs or intentions to others.
Cognitive maps in rats and men
Edward C Tolman · 1948
Earlier work this paper cites.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Bargaining in ignorance of the opponent’s utility function
John C Harsanyi · 1962
Earlier work this paper cites.
Games with incomplete information played by bayesian players, i–iii part i. the basic model
John C Harsanyi · 1967
Earlier work this paper cites.
The optimal control of partially observable markov processes
Edward Jay Sondik · 1971
Earlier work this paper cites.
Does the chimpanzee have a theory of mind?
David Premack and Guy Woodruff · 1978
Earlier work this paper cites.
Folk psychology as simulation
Robert M Gordon · 1986
Earlier work this paper cites.
Two contrasts: folk craft versus folk science, and belief versus opinion
Daniel C Dennett · 1991
Earlier work this paper cites.
A decision-theoretic approach to coordinating multi-agent interactions
Piotr J Gmytrasiewicz, Edmund H Durfee, and David K Wehe · 1991
Earlier work this paper cites.
Why the child’s theory of mind really is a theory
Alison Gopnik and Henry M Wellman · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
A rigorous, operational formalization of recursive modeling
Piotr J Gmytrasiewicz and Edmund H Durfee · 1995
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Mirror neurons and the simulation theory of mind-reading
Vittorio Gallese and Alvin Goldman · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
An introduction to variational methods for graphical models
Michael I Jordan, Zoubin Ghahramani, Tommi S Jaakkola, and Lawrence K Saul · 1999
Earlier work this paper cites.
Rational coordination in multi-agent environments
Piotr J Gmytrasiewicz and Edmund H Durfee · 2000
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games
Satinder Singh, Michael Kearns, and Yishay Mansour · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Friend-or-foe q-learning in general-sum games
Michael L Littman · 2001
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Michael Bowling and Manuela Veloso · 2002
Earlier work this paper cites.
A language for modeling agents’ decision making processes in games
Ya’akov Gal and Avi Pfeffer · 2003
Earlier work this paper cites.
Correlated q-learning
Amy Greenwald, Keith Hall, and Roberto Serrano · 2003
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Cited alongside, same era.
A cognitive hierarchy model of games
Colin F Camerer, Teck-Hua Ho, and Juin-Kuan Chong · 2004
Cited alongside, same era.
Convergence and no-regret in multiagent learning
Michael Bowling · 2005
Cited alongside, same era.
A framework for sequential planning in multi-agent settings
Piotr J Gmytrasiewicz and Prashant Doshi · 2005
Cited alongside, same era.
Psychsim: Modeling theory of mind with decision-theoretic agents
David V Pynadath and Stacy C Marsella · 2005
Cited alongside, same era.
Dealing with non-stationary environments using context detection
Bruno C Da Silva, Eduardo W Basso, Ana LC Bazzan, and Paulo M Engel · 2006
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Planning over multi-agent epistemic states: A classical planning approach
Christian J Muise, Vaishak Belle, Paolo Felli, Sheila A McIlraith, Tim Miller, Adrian R Pearce, and Liz Sonenberg · 2015
Later among the works it cites.
Scalable solutions of interactive pomdps using generalized and bounded policy iteration
Ekhlas Sonu and Prashant Doshi · 2015
Later among the works it cites.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Later among the works it cites.
Opponent modeling in deep reinforcement learning
He He, Jordan Boyd-Graber, Kevin Kwok, and Hal Daumé III · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the difficulty of achieving equilibrium in interactive pomdps
Prashant Doshi and Piotr J Gmytrasiewicz · 2006
Cited alongside, same era.
Biasing coevolutionary search for optimal multiagent behaviors
Liviu Panait, Sean Luke, and R Paul Wiegand · 2006
Cited alongside, same era.
Reaching pareto-optimality in prisoner’s dilemma using conditional joint action learning
Dipyaman Banerjee and Sandip Sen · 2007
Cited alongside, same era.
If multi-agent learning is the answer, what is the question?
Yoav Shoham, Rob Powers, Trond Grenager, et al · 2007
Cited alongside, same era.
Generalized point based value iteration for interactive pomdps
Prashant Doshi and Dennis Perez · 2008
Cited alongside, same era.
Networks of influence diagrams: a formalism for representing agents’ beliefs and decision-making processes
Ya’akov Gal and Avi Pfeffer · 2008
Cited alongside, same era.
Qiang Liu and Dilin Wang · 2016
Later among the works it cites.
Learning to draw samples: With application to amortized mle for generative adversarial learning
Dilin Wang and Qiang Liu · 2016
Later among the works it cites.
Lenient learning in independent-learner stochastic cooperative games
Ermo Wei and Sean Luke · 2016
Later among the works it cites.
Negotiating with other minds: the role of recursive theory of mind in negotiation with incomplete information
Harmen de Weerd, Rineke Verbrugge, and Bart Verheij · 2017
Later among the works it cites.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Later among the works it cites.
The numerics of gans
Lars Mescheder, Sebastian Nowozin, and Andreas Geiger · 2017
Later among the works it cites.
Peng Peng, Ying Wen, Yaodong Yang, Quan Yuan, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Later among the works it cites.
The minds of many: opponent modelling in a stochastic game
Friedrich Burkhard Von Der Osten, Michael Kirley, and Tim Miller · 2017
Later among the works it cites.
Yingce Xia, Tao Qin, Wei Chen, Jiang Bian, Nenghai Yu, and Tie-Yan Liu · 2017
Later among the works it cites.
Autonomous agents modelling other agents: A comprehensive survey and open problems
Stefano V Albrecht and Peter Stone · 2018
Later among the works it cites.
The mechanics of n-player differentiable games
David Balduzzi, Sebastien Racaniere, James Martens, Jakob Foerster, Karl Tuyls, and Thore Graepel · 2018
Later among the works it cites.
Learning with opponent-learning awareness
Jakob Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Later among the works it cites.
Neil C Rabinowitz, Frank Perbet, H Francis Song, Chiyuan Zhang, SM Eslami, and Matthew Botvinick · 2018
Later among the works it cites.
Multiagent soft q-learning
Ermo Wei, Drew Wicke, David Freelan, and Sean Luke · 2018
Later among the works it cites.
Mean field multi-agent reinforcement learning
Yaodong Yang, Rui Luo, Minne Li, Ming Zhou, Weinan Zhang, and Jun Wang · 2018
Later among the works it cites.