2020

The MAGICAL Benchmark for Robust Imitation

Toyer, Sam, Shah, Rohin, Critch, Andrew et al.

Understand

Imitation Learning (IL) algorithms are typically evaluated in the same environment that was used to create demonstrations.

  • This rewards precise reproduction of demonstrations in one particular environment, but provides little information about how robustly an algorithm can generalise the demonstrator's intent to substantially different deployment settings.
  • This paper presents the MAGICAL benchmark suite, which permits systematic evaluation of generalisation by quantifying robustness to different kinds of distribution shift that an IL algorithm is likely to encounter in practice.
  • Using the MAGICAL suite, we confirm that existing IL algorithms overfit significantly to the context in which demonstrations are provided.

Reading the bibliography…