Fetching the paper…
Reading the bibliography…
We study offline multi-agent reinforcement learning (RL) in Markov games, where the goal is to learn an approximate equilibrium -- such as Nash equilibrium and (Coarse) Correlated Equilibrium -- from an offline dataset pre-collected from the game.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Subjectivity and correlation in randomized strategies
Robert J Aumann · 1974
Earlier work this paper cites.
Existence of stationary correlated equilibria with symmetric information for discounted stochastic games
Andrzej S Nowak and TES Raghavan · 1992
Earlier work this paper cites.
Non-cooperative games
John Nash Jr · 1996
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
Online reinforcement learning in stochastic games
Chen-Yu Wei, Yi-Te Hong, and Chi-Jen Lu · 2017
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Earlier work this paper cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Earlier work this paper cites.
Near-optimal reinforcement learning with self-play
Yu Bai, Chi Jin, and Tiancheng Yu · 2020
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with general value function approximation
Ruosong Wang, Ruslan Salakhutdinov, and Lin F Yang · 2020
Cited alongside, same era.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Cited alongside, same era.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Tengyang Xie and Nan Jiang · 2020
Cited alongside, same era.
A sharp analysis of model-based reinforcement learning with self-play
Qinghua Liu, Tiancheng Yu, Yu Bai, and Chi Jin · 2021
Cited alongside, same era.
The complexity of markov equilibrium in stochastic games
Constantinos Daskalakis, Noah Golowich, and Kaiqing Zhang · 2022
Later among the works it cites.
Gap-dependent bounds for two-player markov games
Zehao Dou, Zhuoran Yang, Zhaoran Wang, and Simon Du · 2022
Later among the works it cites.
The power of exploiter: Provable multi-agent rl in large state spaces
Chi Jin, Qinghua Liu, and Tiancheng Yu · 2022
Later among the works it cites.
Settling the sample complexity of model-based offline reinforcement learning
Gen Li, Laixi Shi, Yuxin Chen, Yuejie Chi, and Yuting Wei · 2022
Later among the works it cites.
Provably efficient reinforcement learning in decentralized general-sum markov games
Weichao Mao and Tamer Başar · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Cited alongside, same era.
When can we learn general-sum markov games with a large number of players sample-efficiently?
Ziang Song, Song Mei, and Yu Bai · 2021
Cited alongside, same era.
Pessimistic model-based offline reinforcement learning under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Cited alongside, same era.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2021
Cited alongside, same era.
Towards instance-optimal offline reinforcement learning with pessimism
Ming Yin and Yu-Xiang Wang · 2021
Cited alongside, same era.
Provable benefits of actor-critic methods for offline reinforcement learning
Andrea Zanette, Martin J Wainwright, and Emma Brunskill · 2021
Cited alongside, same era.
Almost optimal algorithms for two-player zero-sum linear mixture markov games
Zixiang Chen, Dongruo Zhou, and Quanquan Gu · 2022
Cited alongside, same era.
Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen, and Yuejie Chi · 2022
Later among the works it cites.
Wei Xiong, Han Zhong, Chengshuai Shi, Cong Shen, Liwei Wang, and Tong Zhang · 2022
Later among the works it cites.
Model-based reinforcement learning is minimax-optimal for offline zero-sum markov games
Yuling Yan, Gen Li, Yuxin Chen, and Jianqing Fan · 2022
Later among the works it cites.
Ming Yin, Yaqi Duan, Mengdi Wang, and Yu-Xiang Wang · 2022
Later among the works it cites.
Offline reinforcement learning with realizability and single-policy concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason Lee · 2022
Later among the works it cites.
Pessimistic minimax value iteration: Provably efficient equilibrium learning from offline datasets
Han Zhong, Wei Xiong, Jiyuan Tan, Liwei Wang, Tong Zhang, Zhaoran Wang, and Zhuoran Yang · 2022
Later among the works it cites.