2023

VDT: General-purpose Video Diffusion Transformers via Mask Modeling

Lu, Haoyu, Yang, Guoxing, Fei, Nanyi et al.

Understand

This work introduces Video Diffusion Transformer (VDT), which pioneers the use of transformers in diffusion-based video generation.

  • It features transformer blocks with modularized temporal and spatial attention modules to leverage the rich spatial-temporal representation inherited in transformers.
  • We also propose a unified spatial-temporal mask modeling mechanism, seamlessly integrated with the model, to cater to diverse video generation scenarios.
  • VDT offers several appealing benefits.

Reading the bibliography…