2018

Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control

Lowrey, Kendall, Rajeswaran, Aravind, Kakade, Sham et al.

Understand

We propose a plan online and learn offline (POLO) framework for the setting where an agent, with an internal model, needs to continually act and learn in the world.

  • Our work builds on the synergistic relationship between local model-based control, global value function learning, and exploration.
  • We study how local trajectory optimization can cope with approximation errors in the value function, and can stabilize and accelerate value function learning.
  • Conversely, we also study how approximate value functions can help reduce the planning horizon and allow for better policies beyond local solutions.

Reading the bibliography…