2023

Learning Universal Policies via Text-Guided Video Generation

Du, Yilun, Yang, Mengjiao, Dai, Bo et al.

Understand

A goal of artificial intelligence is to construct an agent that can solve a wide variety of tasks.

  • Recent progress in text-guided image synthesis has yielded models with an impressive ability to generate complex novel images, exhibiting combinatorial generalization across domains.
  • Motivated by this success, we investigate whether such tools can be used to construct more general-purpose agents.
  • Specifically, we cast the sequential decision making problem as a text-conditioned video generation problem, where, given a text-encoded specification of a desired goal, a planner synthesizes a set of future frames depicting its planned actions in the future, after which control actions are extracted from the generated video.

Reading the bibliography…