Fetching the paper…
Reading the bibliography…
It has long been recognized that multi-agent reinforcement learning (MARL) faces significant scalability issues due to the fact that the size of the state and action spaces are exponentially large in the number of agents.
ALOHA packet system with and without slots and capture
Lawrence Roberts · 1975
Earlier work this paper cites.
Restless bandits: Activity allocation in a changing world
Peter Whittle · 1988
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Neuro-dynamic programming , volume 5
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Solving very large weakly coupled Markov decision processes
Nicolas Meuleau, Milos Hauskrecht, Kee-Eung Kim, Leonid Peshkin, Leslie Pack Kaelbling, Thomas L Dean, and Craig Boutilier · 1998
Earlier work this paper cites.
Efficient reinforcement learning in factored MDPs
Michael Kearns and Daphne Koller · 1999
Earlier work this paper cites.
The complexity of optimal queuing network control
Christos H Papadimitriou and John N Tsitsiklis · 1999
Earlier work this paper cites.
Average cost temporal-difference learning
John N Tsitsiklis and Benjamin Van Roy · 1999
Earlier work this paper cites.
A survey of computational complexity results in systems and control
Vincent D Blondel and John N Tsitsiklis · 2000
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Value-function reinforcement learning in markov games
Michael L Littman · 2001
Earlier work this paper cites.
Distributed control of spatially invariant systems
Bassam Bamieh, Fernando Paganini, and Munther A Dahleh · 2002
Earlier work this paper cites.
Efficient solution algorithms for factored MDPs
Carlos Guestrin, Daphne Koller, Ronald Parr, and Shobha Venkataraman · 2003
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Earlier work this paper cites.
Linear stochastic approximation driven by slowly varying Markov chains
Vijay R Konda and John N Tsitsiklis · 2003
Cited alongside, same era.
Nonequilibrium phase transition in a model for the propagation of innovations among economic agents
Mateu Llas, Pablo M Gleiser, Juan M López, and Albert Díaz-Guilera · 2003
Cited alongside, same era.
The power of epidemics: Robust communication for large-scale distributed systems
Werner Vogels, Robbert van Renesse, and Ken Birman · 2003
Cited alongside, same era.
Dynamic Programming and Optimal Control, Vol. II
Dimitri P. Bertsekas · 2007
Cited alongside, same era.
A comprehensive survey of multiagent reinforcement learning
Lucian Bu, Robert Babu, Bart De Schutter, et al · 2008
Cited alongside, same era.
Epidemic thresholds in real networks
Deepayan Chakrabarti, Yang Wang, Chenxi Wang, Jurij Leskovec, and Christos Faloutsos · 2008
Optimal control of multiroom HVAC system: An event-based approach
Zijian Wu, Qing-Shan Jia, and Xiaohong Guan · 2016
Later among the works it cites.
Control of robotic mobility-on-demand systems: a queueing-theoretical perspective
Rick Zhang and Marco Pavone · 2016
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Later among the works it cites.
Distributed reinforcement learning via gossip
Adwaitvedant Mathkar and Vivek S Borkar · 2017
Later among the works it cites.
On the dynamics of deterministic epidemic propagation over networks
Wenjun Mei, Shadi Mohagheghi, Sandro Zampieri, and Francesco Bullo · 2017
Later among the works it cites.
Decentralized and distributed temperature control via HVAC systems in energy efficient buildings
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Optimal control of spatially distributed systems
Nader Motee and Ali Jadbabaie · 2008
Cited alongside, same era.
A constructive proof of the general Lovász local lemma
Robin A. Moser and Gábor Tardos · 2010
Cited alongside, same era.
Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
Laetitia Matignon, Guillaume J Laurent, and Nadine Le Fort-Piat · 2012
Cited alongside, same era.
Correlation decay method for decision, optimization, and inference in large-scale networks
David Gamarnik · 2013
Cited alongside, same era.
QD-learning: A collaborative distributed strategy for multi-agent reinforcement learning through consensus + innovations
Soummya Kar, José MF Moura, and H Vincent Poor · 2013
Cited alongside, same era.
Correlation decay in random decision networks
David Gamarnik, David A Goldberg, and Theophane Weber · 2014
Cited alongside, same era.
Xuan Zhang, Wenbo Shi, Bin Yan, Ali Malkawi, and Na Li · 2017
Later among the works it cites.
Multi-agent reinforcement learning via double averaging primal-dual optimization
Hoi-To Wai, Zhuoran Yang, Zhaoran Wang, and Mingyi Hong · 2018
Later among the works it cites.
Fully decentralized multi-agent reinforcement learning with networked agents
Kaiqing Zhang, Zhuoran Yang, Han Liu, Tong Zhang, and Tamer Başar · 2018
Later among the works it cites.
Reinforcement learning and deep learning based lateral control for autonomous driving [application notes]
Dong Li, Dongbin Zhao, Qichao Zhang, and Yaran Chen · 2019
Later among the works it cites.
Exploiting fast decaying and locality in multi-agent MDP with tree dependence structure
Guannan Qu and Na Li · 2019
Later among the works it cites.
Scalable reinforcement learning of localized policies for multi-agent networked systems
Guannan Qu, Adam Wierman, and Na Li · 2019
Later among the works it cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar · 2019
Later among the works it cites.
Temporal starvation in multi-channel csma networks: an analytical framework
Alessandro Zocca · 2019
Later among the works it cites.
Q-learning for mean-field controls, 2020
Haotian Gu, Xin Guo, Xiaoli Wei, and Renyuan Xu · 2020
Closest in time.
A finite time analysis of two time-scale actor critic methods, 2020
Yue Wu, Weitong Zhang, Pan Xu, and Quanquan Gu · 2020
Closest in time.
Non-asymptotic convergence analysis of two time-scale (natural) actor-critic algorithms
Tengyu Xu, Zhe Wang, and Yingbin Liang · 2020
Closest in time.