2021

Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning

Agarwal, Rishabh, Machado, Marlos C., Castro, Pablo Samuel et al.

Understand

Reinforcement learning methods trained on few environments rarely learn policies that generalize to unseen environments.

  • To improve generalization, we incorporate the inherent sequential structure in reinforcement learning into the representation learning process.
  • This approach is orthogonal to recent approaches, which rarely exploit this structure explicitly.
  • Specifically, we introduce a theoretically motivated policy similarity metric (PSM) for measuring behavioral similarity between states.

Reading the bibliography…