2022

Diffusion Models for Video Prediction and Infilling

Höppe, Tobias, Mehrjou, Arash, Bauer, Stefan et al.

Understand

Predicting and anticipating future outcomes or reasoning about missing information in a sequence are critical skills for agents to be able to make intelligent decisions.

  • This requires strong, temporally coherent generative capabilities.
  • Diffusion models have shown remarkable success in several generative tasks, but have not been extensively explored in the video domain.
  • We present Random-Mask Video Diffusion (RaMViD), which extends image diffusion models to videos using 3D convolutions, and introduces a new conditioning technique during training.

Reading the bibliography…