2021

Almost Optimal Algorithms for Two-player Zero-Sum Linear Mixture Markov Games

Chen, Zixiang, Zhou, Dongruo, Gu, Quanquan

Understand

We study reinforcement learning for two-player zero-sum Markov games with simultaneous moves in the finite-horizon setting, where the transition kernel of the underlying Markov games can be parameterized by a linear function over the current state, both players' actions and the next state.

  • In particular, we assume that we can control both players and aim to find the Nash Equilibrium by minimizing the duality gap.
  • We propose an algorithm Nash-UCRL based on the principle "Optimism-in-Face-of-Uncertainty".
  • Our algorithm only needs to find a Coarse Correlated Equilibrium (CCE), which is computationally efficient.

Reading the bibliography…