Fetching the paper…
Reading the bibliography…
Though limited in real-world decision making, most multi-agent reinforcement learning (MARL) models assume perfectly rational agents -- a property hardly met due to individual's cognitive limitation and/or the tractability of the decision problem.
Equilibrium points in n-person games
John F Nash et al · 1950
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Theories of bounded rationality
Herbert A Simon · 1972
Earlier work this paper cites.
The General Theory of Employment, Interest and Money
J. M. Keynes · 1973
Earlier work this paper cites.
Sequential equilibria
David M Kreps and Robert Wilson · 1982
Earlier work this paper cites.
Quantal response equilibria for normal form games
Richard D McKelvey and Thomas R Palfrey · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Convergence of gradient dynamics with a variable learning rate
Michael Bowling and Manuela Veloso · 2001
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Michael Bowling and Manuela Veloso · 2002
Earlier work this paper cites.
Games and phone numbers: Do short-term memory bounds affect strategic behavior?
Giovanna Devetag and Massimo Warglien · 2003
Earlier work this paper cites.
Nonlinear control systems: analysis and design
Horacio J Marquez · 2003
Earlier work this paper cites.
Multi-agent reinforcement learning: a critical survey
Yoav Shoham, Rob Powers, and Trond Grenager · 2003
Earlier work this paper cites.
A cognitive hierarchy model of games
Colin F Camerer, Teck-Hua Ho, and Juin-Kuan Chong · 2004
Cited alongside, same era.
A framework for sequential planning in multi-agent settings
Piotr J Gmytrasiewicz and Prashant Doshi · 2005
Cited alongside, same era.
The case for dynamic difficulty adjustment in games
Robin Hunicke · 2005
Cited alongside, same era.
A multiagent reinforcement learning algorithm with non-linear dynamics
Sherief Abdallah and Victor Lesser · 2008
Cited alongside, same era.
Neural correlates of depth of strategic reasoning in medial prefrontal cortex
Giorgio Coricelli and Rosemarie Nagel · 2009
Cited alongside, same era.
Multi-agent learning with policy prediction
Chongjie Zhang and Victor Lesser · 2010
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Later among the works it cites.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games
Peng Peng, Quan Yuan, Ying Wen, Yaodong Yang, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Later among the works it cites.
Autonomous agents modelling other agents: A comprehensive survey and open problems
Stefano V Albrecht and Peter Stone · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Theory of mind
Alvin I Goldman et al · 2012
Cited alongside, same era.
Bounded rationality, abstraction, and hierarchical decision-making: An information-theoretic optimality principle
Tim Genewein, Felix Leibfried, Jordi Grau-Moya, and Daniel Alexander Braun · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Opponent modeling in deep reinforcement learning
He He, Jordan Boyd-Graber, Kevin Kwok, and Hal Daumé III · 2016
Cited alongside, same era.
Stein variational gradient descent: A general purpose bayesian inference algorithm
Qiang Liu and Dilin Wang · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Cited alongside, same era.
Sergey Levine · 2018
Later among the works it cites.
Machine theory of mind
Neil C Rabinowitz, Frank Perbet, H Francis Song, Chiyuan Zhang, SM Eslami, and Matthew Botvinick · 2018
Later among the works it cites.
Lyapunov functions for first-order methods: Tight automated convergence guarantees
Adrien Taylor, Bryan Van Scoy, and Laurent Lessard · 2018
Later among the works it cites.
Multiagent soft q-learning
Ermo Wei, Drew Wicke, David Freelan, and Sean Luke · 2018
Later among the works it cites.
Mean field multi-agent reinforcement learning
Yaodong Yang, Rui Luo, Minne Li, Ming Zhou, Weinan Zhang, and Jun Wang · 2018
Later among the works it cites.
Bridging level-k to nash equilibrium
Dan Levin and Luyao Zhang · 2019
Closest in time.
Theory of minds: Understanding behavior in groups through inverse planning
Michael Shum, Max Kleiman-Weiner, Michael L Littman, and Joshua B Tenenbaum · 2019
Closest in time.
A regularized opponent model with maximum entropy objective
Zheng Tian, Ying Wen, Zhicheng Gong, Faiz Punakkath, Shihao Zou, and Jun Wang · 2019
Closest in time.
Probabilistic recursive reasoning for multi-agent reinforcement learning
Ying Wen, Yaodong Yang, Rui Luo, Jun Wang, and Wei Pan · 2019
Closest in time.