Fetching the paper…
Reading the bibliography…
Despite the increasing interest in multi-agent reinforcement learning (MARL) in multiple communities, understanding its theoretical foundation has long been recognized as a challenging problem.
A theoretical analysis of deep Q-learning
Yang, Z · 1901
Earlier work this paper cites.
A multi-agent off-policy actor-critic algorithm for distributed reinforcement learning
Suttle, W · 1903
Earlier work this paper cites.
Doan, T. T · 1907
Earlier work this paper cites.
Stochastic games
Shapley, L. S · 1953
Earlier work this paper cites.
On stochastic games, i
Maitra, A · 1970
Earlier work this paper cites.
On stochastic games, ii
Maitra, A · 1971
Earlier work this paper cites.
Mixing conditions for Markov chains
Davydov, Y. A · 1973
Earlier work this paper cites.
Discounted, positive, and noncooperative stochastic games
Parthasarathy, T · 1973
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L · 1994
Earlier work this paper cites.
A Course in Game Theory
Osborne, M. J · 1994
Earlier work this paper cites.
Sphere packing numbers for subsets of the boolean n-cube with bounded Vapnik-Chervonenkis dimension
Haussler, D · 1995
Earlier work this paper cites.
Planning, learning and coordination in multi-agent decision processes
Boutilier, C · 1996
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications
Littman, M. L · 1996
Earlier work this paper cites.
RoboCup: A challenge problem for AI
Kitano, H · 1997
Earlier work this paper cites.
Stochastic and shortest path games: Theory and algorithms
Patek, S. D · 1997
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. N · 1997
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
Lauer, M · 2000
Earlier work this paper cites.
Friend-or-Foe Q-learning in general-sum games
Littman, M. L · 2001
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Bowling, M · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M · 2003
Earlier work this paper cites.
Learning in zero-sum team Markov games using factored value functions
Lagoudakis, M. G · 2003
Earlier work this paper cites.
Reinforcement learning to play an optimal Nash equilibrium in team Markov games
Wang, X · 2003
Earlier work this paper cites.
Information flow and cooperative control of vehicle formations
Alexander, F. J · 2004
Earlier work this paper cites.
Networked robots: Flying robot navigation using a sensor net
Corke, P · 2005
Earlier work this paper cites.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Cited alongside, same era.
Cyclic equilibria in Markov games
Zinkevich, M · 2006
Cited alongside, same era.
A comprehensive survey of multi-agent reinforcement learning
Busoniu, L · 2008
Cited alongside, same era.
Finite-time bounds for fitted value iteration
Munos, R · 2008
Cited alongside, same era.
Achieving goals in decentralized POMDPs
Amato, C · 2009
Cited alongside, same era.
Challenges for securing cyber physical systems
Cardenas, A · 2009
Cited alongside, same era.
A Concise Introduction to Decentralized POMDPs
Oliehoek, F. A · 2016
Later among the works it cites.
Learning Nash equilibrium for general-sum Markov games from batch data
Perolat, J · 2016
Later among the works it cites.
Do Nascimento Silva, V · 2017
Later among the works it cites.
Stochastic proximal gradient consensus over random networks
Hong, M · 2017
Later among the works it cites.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning in finite MDPs: PAC analysis
Strehl, A. L · 2009
Cited alongside, same era.
Finite-sample analysis of LSTD
Lazaric, A · 2010
Cited alongside, same era.
Sample bounded distributed reinforcement learning for decentralized POMDPs
Banerjee, B · 2012
Cited alongside, same era.
Mixing: Properties and Examples
Doukhan, P · 2012
Cited alongside, same era.
A tutorial on linear function approximators for dynamic programming and reinforcement learning
Geramifard, A · 2013
Cited alongside, same era.
QD-learning: A collaborative distributed strategy for multi-agent reinforcement learning through Consensus + Innovations
Kar, S · 2013
Cited alongside, same era.
Lowe, R · 2017
Later among the works it cites.
Distributed reinforcement learning via gossip
Mathkar, A · 2017
Later among the works it cites.
Achieving geometric convergence for distributed optimization over time-varying graphs
Nedic, A · 2017
Later among the works it cites.
The duality gap for two-team zero-sum games
Schulman, L · 2017
Later among the works it cites.
Multiagent cooperation and competition with deep reinforcement learning
Tampuu, A · 2017
Later among the works it cites.
Non-convex distributed optimization
Tatarenko, T · 2017
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J · 2018
Closest in time.
Communication-efficient distributed reinforcement learning
Chen, T · 2018
Closest in time.
Finite sample analyses for TD(0) with function approximation
Dalal, G · 2018
Closest in time.
An incremental off-policy search in a model-free Markov decision process using a single sample path
Joseph, A. G · 2018
Closest in time.
Primal-dual algorithm for distributed reinforcement learning: Distributed GTD2
Lee, D · 2018
Closest in time.
Actor-critic policy optimization in partially observable multiagent environments
Srinivasan, S · 2018
Closest in time.
Multi-agent reinforcement learning via double averaging primal-dual optimization
Wai, H.-T · 2018
Closest in time.
A finite sample analysis of the actor-critic algorithm
Yang, Z · 2018
Closest in time.
Information-theoretic considerations in batch reinforcement learning
Chen, J · 2019
Closest in time.
Finite-time error bounds for linear stochastic approximation and td learning
Srikant, R · 2019
Closest in time.
Policy optimization provably converges to Nash equilibria in zero-sum linear quadratic games
Zhang, K · 2019
Closest in time.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Zhang, K · 2020
Closest in time.