2022

MAGVIT: Masked Generative Video Transformer

Yu, Lijun, Cheng, Yong, Sohn, Kihyuk et al.

Understand

We introduce the MAsked Generative VIdeo Transformer, MAGVIT, to tackle various video synthesis tasks with a single model.

  • We introduce a 3D tokenizer to quantize a video into spatial-temporal visual tokens and propose an embedding method for masked video token modeling to facilitate multi-task learning.
  • We conduct extensive experiments to demonstrate the quality, efficiency, and flexibility of MAGVIT.
  • Our experiments show that (i) MAGVIT performs favorably against state-of-the-art approaches and establishes the best-published FVD on three video generation benchmarks, including the challenging Kinetics-600.

Reading the bibliography…