2023

Photorealistic Video Generation with Diffusion Models

Gupta, Agrim, Yu, Lijun, Sohn, Kihyuk et al.

Understand

We present W.A.L.T, a transformer-based approach for photorealistic video generation via diffusion modeling.

  • Our approach has two key design decisions.
  • First, we use a causal encoder to jointly compress images and videos within a unified latent space, enabling training and generation across modalities.
  • Second, for memory and training efficiency, we use a window attention architecture tailored for joint spatial and spatiotemporal generative modeling.

Reading the bibliography…