2021

Can Reinforcement Learning Find Stackelberg-Nash Equilibria in General-Sum Markov Games with Myopic Followers?

Zhong, Han, Yang, Zhuoran, Wang, Zhaoran et al.

Understand

We study multi-player general-sum Markov games with one of the players designated as the leader and the other players regarded as followers.

  • In particular, we focus on the class of games where the followers are myopic, i.e., they aim to maximize their instantaneous rewards.
  • For such a game, our goal is to find a Stackelberg-Nash equilibrium (SNE), which is a policy pair $(\pi^*, \nu^*)$ such that (i) $\pi^*$ is the optimal policy for the leader when the followers always play their best response, and (ii) $\nu^*$ is the best response policy of the followers, which is a Nash equilibrium of the followers' game induced by $\pi^*$.
  • We develop sample-efficient reinforcement learning (RL) algorithms for solving for an SNE in both online and offline settings.

Reading the bibliography…