Fetching the paper…
Reading the bibliography…
Markov Games (MG) is an important model for Multi-Agent Reinforcement Learning (MARL).
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
On tail probabilities for martingales
David A Freedman · 1975
Earlier work this paper cites.
The complexity of computing a nash equilibrium
Constantinos Daskalakis, Paul W Goldberg, and Christos H Papadimitriou · 2009
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Efficient reinforcement learning in deterministic systems with value function generalization
Zheng Wen and Benjamin Van Roy · 2017
Earlier work this paper cites.
Superhuman ai for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Earlier work this paper cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Earlier work this paper cites.
Near-optimal reinforcement learning with self-play
Yu Bai, Chi Jin, and Tiancheng Yu · 2020
Earlier work this paper cites.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2020
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Earlier work this paper cites.
Bias no more: high-probability data-dependent regret bounds for adversarial bandits and mdps
Chung-Wei Lee, Haipeng Luo, Chen-Yu Wei, and Mengxiao Zhang · 2020
Cited alongside, same era.
Efficient and robust algorithms for adversarial linear contextual bandits
Gergely Neu and Julia Olkhovskaya · 2020
Cited alongside, same era.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Cited alongside, same era.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Cited alongside, same era.
V-learning–a simple, efficient, decentralized algorithm for multiagent rl
Chi Jin, Qinghua Liu, Yuanhao Wang, and Tiancheng Yu · 2021
Cited alongside, same era.
Stabilizing q-learning with linear architectures for provable efficient learning
Andrea Zanette and Martin Wainwright · 2022
Later among the works it cites.
Return of the bias: Almost minimax optimal high probability bounds for adversarial linear bandits
Julian Zimmert and Tor Lattimore · 2022
Later among the works it cites.
Breaking the curse of multiagents in a large state space: Rl in markov games with independent linear function approximation
Qiwen Cui, Kaiqing Zhang, and Simon Du · 2023
Later among the works it cites.
Refined regret for adversarial mdps with linear function approximation
Yan Dai, Haipeng Luo, Chen-Yu Wei, and Julian Zimmert · 2023
Later among the works it cites.
The complexity of markov equilibrium in stochastic games
Constantinos Daskalakis, Noah Golowich, and Kaiqing Zhang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A sharp analysis of model-based reinforcement learning with self-play
Qinghua Liu, Tiancheng Yu, Yu Bai, and Chi Jin · 2021
Cited alongside, same era.
Policy optimization in adversarial mdps: Improved exploration via dilated bonuses
Haipeng Luo, Chen-Yu Wei, and Chung-Wei Lee · 2021
Cited alongside, same era.
Towards general function approximation in zero-sum markov games
Baihe Huang, Jason D. Lee, Zhaoran Wang, and Zhuoran Yang · 2022
Cited alongside, same era.
The power of exploiter: Provable multi-agent rl in large state spaces
Chi Jin, Qinghua Liu, and Tiancheng Yu · 2022
Cited alongside, same era.
Representation learning for general-sum low-rank markov games
Chengzhuo Ni, Yuda Song, Xuezhou Zhang, Chi Jin, and Mengdi Wang · 2022
Cited alongside, same era.
When can we learn general-sum markov games with a large number of players sample-efficiently?
Ziang Song, Song Mei, and Yu Bai · 2022
Cited alongside, same era.
A self-play posterior sampling algorithm for zero-sum markov games
Wei Xiong, Han Zhong, Chengshuai Shi, Cong Shen, and Tong Zhang · 2022
Cited alongside, same era.
Fang Kong, Xiangcheng Zhang, Baoxiang Wang, and Shuai Li · 2023
Later among the works it cites.
Provably efficient reinforcement learning in decentralized general-sum markov games
Weichao Mao and Tamer Başar · 2023
Later among the works it cites.
First-and second-order bounds for adversarial linear contextual bandits
Julia Olkhovskaya, Jack Mayo, Tim van Erven, Gergely Neu, and Chen-Yu Wei · 2023
Later among the works it cites.
Breaking the curse of multiagency: Provably efficient decentralized multi-agent rl with function approximation
Yuanhao Wang, Qinghua Liu, Yu Bai, and Chi Jin · 2023
Later among the works it cites.
Decentralized optimistic hyperpolicy mirror descent: Provably no-regret learning in markov games
Wenhao Zhan, Jason D. Lee, and Zhuoran Yang · 2023
Later among the works it cites.
Learning adversarial linear mixture markov decision processes with bandit feedback and unknown transition
Canzhe Zhao, Ruofeng Yang, Baoxiang Wang, and Shuai Li · 2023
Later among the works it cites.
Rl in markov games with independent function approximation: Improved sample complexity bound under the local access model
Junyi Fan, Yuxuan Han, Jialin Zeng, Jian-Feng Cai, Yang Wang, Yang Xiang, and Jiheng Zhang · 2024
Closest in time.