2021

Behavioral Priors and Dynamics Models: Improving Performance and Domain Transfer in Offline RL

Cang, Catherine, Rajeswaran, Aravind, Abbeel, Pieter et al.

Understand

Offline Reinforcement Learning (RL) aims to extract near-optimal policies from imperfect offline data without additional environment interactions.

  • Extracting policies from diverse offline datasets has the potential to expand the range of applicability of RL by making the training process safer, faster, and more streamlined.
  • We investigate how to improve the performance of offline RL algorithms, its robustness to the quality of offline data, as well as its generalization capabilities.
  • To this end, we introduce Offline Model-based RL with Adaptive Behavioral Priors (MABE).

Reading the bibliography…