2019

Imagined Value Gradients: Model-Based Policy Optimization with Transferable Latent Dynamics Models

Byravan, Arunkumar, Springenberg, Jost Tobias, Abdolmaleki, Abbas et al.

Understand

Humans are masters at quickly learning many complex tasks, relying on an approximate understanding of the dynamics of their environments.

  • In much the same way, we would like our learning agents to quickly adapt to new tasks.
  • In this paper, we explore how model-based Reinforcement Learning (RL) can facilitate transfer to new tasks.
  • We develop an algorithm that learns an action-conditional, predictive model of expected future observations, rewards and values from which a policy can be derived by following the gradient of the estimated value along imagined trajectories.

Reading the bibliography…