Fetching the paper…
Reading the bibliography…
Multi-agent reinforcement learning (MARL) is often modeled using the framework of Markov games (also called stochastic games or dynamic games).
Stochastic games
L. S. Shapley · 1953
Earlier work this paper cites.
Equilibrium in a stochastic $n$-person game
A. M. Fink · 1964
Earlier work this paper cites.
Equilibrium points of stochastic non-cooperative n n -person games
M. Takahashi · 1964
Earlier work this paper cites.
On nonterminating stochastic games
A. J. Hoffman and R. M. Karp · 1966
Earlier work this paper cites.
Nonzero-sum stochastic games
P. D. Rogers · 1969
Earlier work this paper cites.
A birth–death model of advertising and pricing
S. C. Albright and W. Winston · 1979
Earlier work this paper cites.
Representation and approximation of noncooperative sequential games
W. Whitt · 1980
Earlier work this paper cites.
Stochastic games with finite state and action spaces
O. J. Vrieze · 1987
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
Algorithms for Stochastic Games , pages 45–57
M. Breton · 1991
Earlier work this paper cites.
Nonlinear programming and stationary equilibria in stochastic games
J. A. Filar, T. A. Schultz, F. Thuijsman, and O. Vrieze · 1991
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Markov-perfect industry dynamics: A framework for empirical work
R. Ericson and A. Pakes · 1995
Earlier work this paper cites.
Competitive Markov Decision Processes
J. Filar and K. Vrieze · 1996
Earlier work this paper cites.
Approximations in dynamic zero-sum games I
M. M. Tidball and E. Altman · 1996
Earlier work this paper cites.
How does the value function of a Markov decision process depend on the transition probabilities?
A. Müller · 1997
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
A. Müller · 1997
Earlier work this paper cites.
Approximations in dynamic zero-sum games II
M. M. Tidball, O. Pourtallier, and E. Altman · 1997
Earlier work this paper cites.
Constrained Markov decision processes: stochastic modeling
E. Altman · 1999
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
M. Kearns and S. Singh · 1999
Earlier work this paper cites.
A dynamic oligopoly with collusion and price wars
C. Fershtiman and A. Pakes · 2000
Cited alongside, same era.
A theory of political transitions
D. Acemoglu and J. A. Robinson · 2001
Cited alongside, same era.
Value-function reinforcement learning in Markov games
M. L. Littman · 2001
Cited alongside, same era.
Markov perfect equilibrium: I. observable actions
E. Maskin and J. Tirole · 2001
Cited alongside, same era.
On the sample complexity of reinforcement learning
S. M. Kakade · 2003
Cited alongside, same era.
Multi-agent reinforcement learning: a critical survey
Y. Shoham, R. Powers, and T. Grenager · 2003
Cited alongside, same era.
Stationary equilibria in stochastic games: Structure, selection, and computation
Homotopy methods to compute equilibria in game theory
P. J.-J. Herings and R. Peeters · 2010
Later among the works it cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
M. G. Azar, R. Munos, and H. J. Kappen · 2013
Later among the works it cites.
Robust Markov perfect equilibria
A. Jaśkiewicz and A. S. Nowak · 2014
Later among the works it cites.
Two-timescale algorithms for learning Nash equilibria in general-sum stochastic games
H. Prasad, P. LA, and S. Bhatnagar · 2015
Later among the works it cites.
Dynamic programming and optimal control
D. P. Bertsekas · 2017
Later among the works it cites.
Learning Nash equilibrium for general-sum Markov games from batch data
J. Pérolat, F. Strub, B. Piot, and O. Pietquin · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. J.-J. Herings, R. J. Peeters, et al · 2004
Cited alongside, same era.
Lipschitz continuity of value functions in Markovian decision processes
K. Hinderer · 2005
Cited alongside, same era.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Cited alongside, same era.
Repeated Games and Reputations: Long-Run Relationships
G. J. Mailath and L. Samuelson · 2006
Cited alongside, same era.
Cyclic equilibria in Markov games
M. Zinkevich, A. Greenwald, and M. Littman · 2006
Cited alongside, same era.
Sequential estimation of dynamic discrete games
V. Aguirregabiria and P. Mira · 2007
Cited alongside, same era.
Handbook of Dynamic Game Theory
T. Başar and G. Zaccour, editors · 2018
Later among the works it cites.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
A. Sidford, M. Wang, X. Wu, L. F. Yang, and Y. Ye · 2018
Later among the works it cites.
Multi-agent reinforcement learning with multi-step generative models
O. Krupnik, I. Mordatch, and A. Tamar · 2019
Later among the works it cites.
General sum markov games for strategic detection of advanced persistent threats using moving target defense in cloud networks
S. Sengupta, A. Chowdhary, D. Huang, and S. Kambhampati · 2019
Later among the works it cites.
Benchmarking Model-Based Reinforcement Learning
T. Wang, X. Bao, I. Clavera, J. Hoang, Y. Wen, E. Langlois, S. Zhang, G. Zhang, P. Abbeel, and J. Ba · 2019
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
A. Agarwal, S. Kakade, and L. F. Yang · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
G. Li, Y. Wei, Y. Chi, Y. Gu, and Y. Chen · 2020
Later among the works it cites.
Solving discounted stochastic two-player games with near-optimal time and sample complexity
A. Sidford, M. Wang, L. Yang, and Y. Ye · 2020
Later among the works it cites.
Model-based multi-agent rl in zero-sum Markov games with near-optimal sample complexity
K. Zhang, S. Kakade, T. Basar, and L. Yang · 2020
Later among the works it cites.
On the complexity of computing markov perfect equilibrium in general-sum stochastic games
X. Deng, Y. Li, D. H. Mguni, J. Wang, and Y. Yang · 2021
Closest in time.
Global convergence of multi-agent policy gradient in markov potential games
S. Leonardos, W. Overman, I. Panageas, and G. Piliouras · 2021
Closest in time.
A Course in Stochastic Game Theory
E. Solan · 2021
Closest in time.
When can we learn general-sum markov games with a large number of players sample-efficiently?
Z. Song, S. Mei, and Y. Bai · 2021
Closest in time.
Approximate information state for approximate planning and reinforcement learning in partially observed systems
J. Subramanian, A. Sinha, R. Seraj, and A. Mahajan · 2022
Closest in time.