2022

Phenaki: Variable Length Video Generation From Open Domain Textual Description

Villegas, Ruben, Babaeizadeh, Mohammad, Kindermans, Pieter-Jan et al.

Understand

We present Phenaki, a model capable of realistic video synthesis, given a sequence of textual prompts.

  • Generating videos from text is particularly challenging due to the computational cost, limited quantities of high quality text-video data and variable length of videos.
  • To address these issues, we introduce a new model for learning video representation which compresses the video to a small representation of discrete tokens.
  • This tokenizer uses causal attention in time, which allows it to work with variable-length videos.

Reading the bibliography…