2021

Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability

Ghosh, Dibya, Rahme, Jad, Kumar, Aviral et al.

Understand

Generalization is a central challenge for the deployment of reinforcement learning (RL) systems in the real world.

  • In this paper, we show that the sequential structure of the RL problem necessitates new approaches to generalization beyond the well-studied techniques used in supervised learning.
  • While supervised learning methods can generalize effectively without explicitly accounting for epistemic uncertainty, we show that, perhaps surprisingly, this is not the case in RL.
  • We show that generalization to unseen test conditions from a limited number of training conditions induces implicit partial observability, effectively turning even fully-observed MDPs into POMDPs.

Reading the bibliography…