2021

DisCo RL: Distribution-Conditioned Reinforcement Learning for General-Purpose Policies

Nasiriany, Soroush, Pong, Vitchyr H., Nair, Ashvin et al.

Understand

Can we use reinforcement learning to learn general-purpose policies that can perform a wide range of different tasks, resulting in flexible and reusable skills? Contextual policies provide this capability in principle, but the representation of the context determines the degree of generalization and expressivity.

  • Categorical contexts preclude generalization to entirely new tasks.
  • Goal-conditioned policies may enable some generalization, but cannot capture all tasks that might be desired.
  • In this paper, we propose goal distributions as a general and broadly applicable task representation suitable for contextual policies.

Reading the bibliography…