Fetching the paper…
Reading the bibliography…
This paper studies a class of multi-agent reinforcement learning (MARL) problems where the reward that an agent receives depends on the states of other agents, but the next state only depends on the agent's own current state and action.
Aloha packet system with and without slots and capture
Lawrence G. Roberts · 1975
Earlier work this paper cites.
Solving very large weakly coupled markov decision processes
Nicolas Meuleau, Milos Hauskrecht, Kee-Eung Kim, Leonid Peshkin, Leslie Pack Kaelbling, Thomas Dean, and Craig Boutilier · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis · 2000
Earlier work this paper cites.
CDMA uplink power control as a non-cooperative game
Tansu Alpcan, Tamer Başar, R. Srikant, and Eitan Altman · 2001
Earlier work this paper cites.
Power control in wireless cellular networks
Mung Chiang, Prashanth Hande, Tian Lan, and Chee Wei Tan · 2008
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Earlier work this paper cites.
Online learning in weakly coupled markov decision processes: A convergence time study
Xiaohan Wei, Hao Yu, and Michael J. Neely · 2018
Earlier work this paper cites.
Fully decentralized multi-agent reinforcement learning with networked agents
Kaiqing Zhang, Zhuoran Yang, Han Liu, Tong Zhang, and Tamer Basar · 2018
Cited alongside, same era.
Learning mean-field games
Xin Guo, Anran Hu, Renyuan Xu, and Junzi Zhang · 2019
Cited alongside, same era.
Finite-time error bounds for linear stochastic approximation and TD learning
R. Srikant and Lei Ying · 2019
Cited alongside, same era.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Cited alongside, same era.
Multi-agent reinforcement learning for networked system control
Tianshu Chu, Sandeep Chinchali, and Sachin Katti · 2020
Cited alongside, same era.
On the convergence of model free learning in mean field games
Romuald Elie, Pérolat Julien, Mathieu Laurière, Matthieu Geist, and Olivier Pietquin · 2020
A universal transcoding and transmission method for livecast with networked multi-agent reinforcement learning
Xingyan Chen, Changqiao Xu, Mu Wang, Zhonghui Wu, Shujie Yang, Lujie Zhong, and Gabriel-Miro Muntean · 2021
Later among the works it cites.
Multi-agent reinforcement learning in stochastic networked systems
Yiheng Lin, Guannan Qu, Longbo Huang, and Adam Wierman · 2021
Later among the works it cites.
MAMRL: Exploiting multi-agent meta reinforcement learning in WAN traffic engineering
Shan Sun, Mariam Kiran, and Wei Ren · 2021
Later among the works it cites.
Learning while playing in mean-field games: Convergence and optimality
Qiaomin Xie, Zhuoran Yang, Zhaoran Wang, and Andreea Minca · 2021
Later among the works it cites.
Sample efficient reinforcement learning with reinforce
Junzi Zhang, Jongho Kim, Brendan O’Donoghue, and Stephen Boyd · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Cited alongside, same era.
Scalable multi-agent reinforcement learning for networked systems with average reward
Guannan Qu, Yiheng Lin, Adam Wierman, and Na Li · 2020
Cited alongside, same era.
Scalable reinforcement learning of localized policies for multi-agent networked systems
Guannan Qu, Adam Wierman, and Na Li · 2020
Cited alongside, same era.
Intelligent video caching at network edge: A multi-agent deep reinforcement learning approach
Fangxin Wang, Feng Wang, Jiangchuan Liu, Ryan Shea, and Lifeng Sun · 2020
Cited alongside, same era.
Runyu Zhang, Zhaolin Ren, and Na Li · 2021
Later among the works it cites.
Sample and communication-efficient decentralized actor-critic algorithms with finite-time analysis
Ziyi Chen, Yi Zhou, Rong-Rong Chen, and Shaofeng Zou · 2022
Closest in time.
Finite-time convergence and sample complexity of multi-agent actor-critic reinforcement learning with average reward
FNU Hairi, Jia Liu, and Songtao Lu · 2022
Closest in time.
Scalable reinforcement learning for multiagent networked systems
Guannan Qu, Adam Wierman, and Na Li · 2022
Closest in time.
Global convergence of localized policy iteration in networked multi-agent reinforcement learning
Yizhou Zhang, Guannan Qu, Pan Xu, Yiheng Lin, Zaiwei Chen, and Adam Wierman · 2022
Closest in time.