2015

From Pixels to Torques: Policy Learning with Deep Dynamical Models

Wahlström, Niklas, Schön, Thomas B., Deisenroth, Marc Peter

Understand

Data-efficient learning in continuous state-action spaces using very high-dimensional observations remains a key challenge in developing fully autonomous systems.

  • In this paper, we consider one instance of this challenge, the pixels to torques problem, where an agent must learn a closed-loop control policy from pixel information only.
  • We introduce a data-efficient, model-based reinforcement learning algorithm that learns such a closed-loop policy directly from pixel information.
  • The key ingredient is a deep dynamical model that uses deep auto-encoders to learn a low-dimensional embedding of images jointly with a predictive model in this low-dimensional feature space.

Reading the bibliography…