2019

BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement Learning

Chen, Xinyue, Zhou, Zijian, Wang, Zheng et al.

Understand

There has recently been a surge in research in batch Deep Reinforcement Learning (DRL), which aims for learning a high-performing policy from a given dataset without additional interactions with the environment.

  • We propose a new algorithm, Best-Action Imitation Learning (BAIL), which strives for both simplicity and performance.
  • BAIL learns a V function, uses the V function to select actions it believes to be high-performing, and then uses those actions to train a policy network using imitation learning.
  • For the MuJoCo benchmark, we provide a comprehensive experimental study of BAIL, comparing its performance to four other batch Q-learning and imitation-learning schemes for a large variety of batch datasets.

Reading the bibliography…