2021

Asymmetric self-play for automatic goal discovery in robotic manipulation

OpenAI, OpenAI, Plappert, Matthias, Sampedro, Raul et al.

Understand

We train a single, goal-conditioned policy that can solve many robotic manipulation tasks, including tasks with previously unseen goals and objects.

  • We rely on asymmetric self-play for goal discovery, where two agents, Alice and Bob, play a game.
  • Alice is asked to propose challenging goals and Bob aims to solve them.
  • We show that this method can discover highly diverse and complex goals without any human priors.

Reading the bibliography…