2023

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Blattmann, Andreas, Dockhorn, Tim, Kulal, Sumith et al.

Understand

We present Stable Video Diffusion - a latent video diffusion model for high-resolution, state-of-the-art text-to-video and image-to-video generation.

  • Recently, latent diffusion models trained for 2D image synthesis have been turned into generative video models by inserting temporal layers and finetuning them on small, high-quality video datasets.
  • However, training methods in the literature vary widely, and the field has yet to agree on a unified strategy for curating video data.
  • In this paper, we identify and evaluate three different stages for successful training of video LDMs: text-to-image pretraining, video pretraining, and high-quality video finetuning.

Reading the bibliography…