2020

"Other-Play" for Zero-Shot Coordination

Hu, Hengyuan, Lerer, Adam, Peysakhovich, Alex et al.

Understand

We consider the problem of zero-shot coordination - constructing AI agents that can coordinate with novel partners they have not seen before (e.g.

  • humans).
  • Standard Multi-Agent Reinforcement Learning (MARL) methods typically focus on the self-play (SP) setting where agents construct strategies by playing the game with themselves repeatedly.
  • Unfortunately, applying SP naively to the zero-shot coordination problem can produce agents that establish highly specialized conventions that do not carry over to novel partners they have not been trained with.

Reading the bibliography…