Fetching the paper…
Reading the bibliography…
We study multi-agent general-sum Markov games with nonlinear function approximation.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Efficient reinforcement learning in factored mdps
Michael Kearns and Daphne Koller · 1999
Earlier work this paper cites.
Algorithm-directed exploration for model-based reinforcement learning in factored mdps
Carlos Guestrin, Relu Patrascu, and Dale Schuurmans · 2002
Earlier work this paper cites.
Efficient solution algorithms for factored mdps
Carlos Guestrin, Daphne Koller, Ronald Parr, and Shobha Venkataraman · 2003
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Earlier work this paper cites.
Efficient structure learning in factored-state mdps
Alexander L Strehl, Carlos Diuk, and Michael L Littman · 2007
Earlier work this paper cites.
Regret minimization and the price of total anarchy
Avrim Blum, MohammadTaghi Hajiaghayi, Katrina Ligett, and Aaron Roth · 2008
Earlier work this paper cites.
Computing correlated equilibria in multi-player games
Christos H Papadimitriou and Tim Roughgarden · 2008
Earlier work this paper cites.
Algorithmic game theory
Tim Roughgarden · 2010
Earlier work this paper cites.
Swarm robotics: a review from the swarm engineering perspective
Manuele Brambilla, Eliseo Ferrante, Mauro Birattari, and Marco Dorigo · 2013
Earlier work this paper cites.
On the complexity of approximating a nash equilibrium
Constantinos Daskalakis · 2013
Earlier work this paper cites.
Strategy iteration is strongly polynomial for 2-player turn-based stochastic games with a constant discount factor
Thomas Dueholm Hansen, Peter Bro Miltersen, and Uri Zwick · 2013
Earlier work this paper cites.
Fictitious self-play in extensive-form games
Johannes Heinrich, Marc Lanctot, and David Silver · 2015
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Exclusion method for finding nash equilibrium in multiplayer games
Kimmo Berg and Tuomas Sandholm · 2017
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2020
Later among the works it cites.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Later among the works it cites.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 2020
Later among the works it cites.
Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity
Kaiqing Zhang, Sham Kakade, Tamer Basar, and Lin Yang · 2020
Later among the works it cites.
Sample-efficient learning of stackelberg equilibria in general-sum games
Yu Bai, Chi Jin, Huan Wang, and Caiming Xiong · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
Provably efficient rl with rich observations via latent state decoding
Simon Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudik, and John Langford · 2019
Cited alongside, same era.
Feature-based q-learning for two-player stochastic games
Zeyu Jia, Lin F Yang, and Mengdi Wang · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Cited alongside, same era.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Cited alongside, same era.
Zixiang Chen, Dongruo Zhou, and Quanquan Gu · 2021
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Simon Du, Sham Kakade, Jason Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Later among the works it cites.
Towards general function approximation in zero-sum markov games
Baihe Huang, Jason D Lee, Zhaoran Wang, and Zhuoran Yang · 2021
Later among the works it cites.
A sharp analysis of model-based reinforcement learning with self-play
Qinghua Liu, Tiancheng Yu, Yu Bai, and Chi Jin · 2021
Later among the works it cites.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Later among the works it cites.
A free lunch from the noise: Provable and practical exploration for representation learning
Tongzheng Ren, Tianjun Zhang, Csaba Szepesvári, and Bo Dai · 2021
Later among the works it cites.
Representation learning for online and offline rl in low-rank mdps
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2021
Later among the works it cites.
Cautiously optimistic policy optimization and exploration with linear function approximation
Andrea Zanette, Ching-An Cheng, and Alekh Agarwal · 2021
Later among the works it cites.
Contrastive ucb: Provably efficient contrastive self-supervised learning in online reinforcement learning
Shuang Qiu, Lingxiao Wang, Chenjia Bai, Zhuoran Yang, and Zhaoran Wang · 2022
Closest in time.
Efficient reinforcement learning in block mdps: A model-free representation learning approach
Xuezhou Zhang, Yuda Song, Masatoshi Uehara, Mengdi Wang, Wen Sun, and Alekh Agarwal · 2022
Closest in time.