Fetching the paper…
Reading the bibliography…
We prove that optimistic-follow-the-regularized-leader (OFTRL), together with smooth value updates, finds an $O(T^{-1})$-approximate Nash equilibrium in $T$ iterations for two-player zero-sum Markov games with full information.
Non-cooperative games
John Nash · 1951
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Earlier work this paper cites.
Excessive gap technique in nonsmooth convex minimization
Yu Nesterov · 2005
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2008
Earlier work this paper cites.
Near-optimal no-regret algorithms for zero-sum games
Constantinos Daskalakis, Alan Deckelbaum, and Anthony Kim · 2011
Earlier work this paper cites.
Optimization, learning, and games with predictable sequences
Sasha Rakhlin and Karthik Sridharan · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Approximate dynamic programming for two-player zero-sum markov games
Julien Perolat, Bruno Scherrer, Bilal Piot, and Olivier Pietquin · 2015
Earlier work this paper cites.
Fast convergence of regularized learning in games
Vasilis Syrgkanis, Alekh Agarwal, Haipeng Luo, and Robert E Schapire · 2015
Earlier work this paper cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Earlier work this paper cites.
Superhuman ai for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Cited alongside, same era.
Feature-based q-learning for two-player stochastic games
Zeyu Jia, Lin F Yang, and Mengdi Wang · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Cited alongside, same era.
Near-optimal reinforcement learning with self-play
Yu Bai, Chi Jin, and Tiancheng Yu · 2020
Cited alongside, same era.
Hedging in games: Faster convergence of external and swap regrets
Provably efficient policy gradient methods for two-player zero-sum markov games
Yulai Zhao, Yuandong Tian, Jason D Lee, and Simon S Du · 2021
Later among the works it cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar · 2021
Later among the works it cites.
Near-optimal no-regret learning for correlated equilibria in multi-player general-sum games
Ioannis Anagnostides, Constantinos Daskalakis, Gabriele Farina, Maxwell Fishelson, Noah Golowich, and Tuomas Sandholm · 2022
Closest in time.
Uncoupled learning dynamics with o ( log t ) o(\log t) swap regret in multiplayer games
Ioannis Anagnostides, Gabriele Farina, Christian Kroer, Chung-Wei Lee, Haipeng Luo, and Tuomas Sandholm · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xi Chen and Binghui Peng · 2020
Cited alongside, same era.
Solving discounted stochastic two-player games with near-optimal time and sample complexity
Aaron Sidford, Mengdi Wang, Lin Yang, and Yinyu Ye · 2020
Cited alongside, same era.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Cited alongside, same era.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yaodong Yang and Jun Wang · 2020
Cited alongside, same era.
Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity
Kaiqing Zhang, Sham Kakade, Tamer Basar, and Lin Yang · 2020
Cited alongside, same era.
Fast policy extragradient methods for competitive games with entropy regularization
Shicong Cen, Yuting Wei, and Yuejie Chi · 2021
Cited alongside, same era.
Near-optimal no-regret learning in general games
Constantinos Daskalakis, Maxwell Fishelson, and Noah Golowich · 2021
Cited alongside, same era.
Ioannis Anagnostides, Ioannis Panageas, Gabriele Farina, and Tuomas Sandholm · 2022
Closest in time.
Faster last-iterate convergence of policy optimization in zero-sum markov games
Shicong Cen, Yuejie Chi, Simon S Du, and Lin Xiao · 2022
Closest in time.
When is offline two-player zero-sum markov game solvable?
Qiwen Cui and Simon S Du · 2022
Closest in time.
Regret minimization and convergence to equilibria in general-sum markov games
Liad Erez, Tal Lancewicki, Uri Sherman, Tomer Koren, and Yishay Mansour · 2022
Closest in time.
Near-optimal no-regret learning for general convex games
Gabriele Farina, Ioannis Anagnostides, Haipeng Luo, Chung-Wei Lee, Christian Kroer, and Tuomas Sandholm · 2022
Closest in time.
Minimax-optimal multi-agent rl in zero-sum markov games with a generative model
Gen Li, Yuejie Chi, Yuting Wei, and Yuxin Chen · 2022
Closest in time.
Model-based reinforcement learning is minimax-optimal for offline zero-sum markov games
Yuling Yan, Gen Li, Yuxin Chen, and Jianqing Fan · 2022
Closest in time.
Policy optimization for markov games: Unified framework and faster convergence
Runyu Zhang, Qinghua Liu, Huan Wang, Caiming Xiong, Na Li, and Yu Bai · 2022
Closest in time.
No-regret learning in time-varying zero-sum games
Mengxiao Zhang, Peng Zhao, Haipeng Luo, and Zhi-Hua Zhou · 2022
Closest in time.