2022

Transframer: Arbitrary Frame Prediction with Generative Models

Nash, Charlie, Carreira, João, Walker, Jacob et al.

Understand

We present a general-purpose framework for image modelling and vision tasks based on probabilistic frame prediction.

  • Our approach unifies a broad range of tasks, from image segmentation, to novel view synthesis and video interpolation.
  • We pair this framework with an architecture we term Transframer, which uses U-Net and Transformer components to condition on annotated context frames, and outputs sequences of sparse, compressed image features.
  • Transframer is the state-of-the-art on a variety of video generation benchmarks, is competitive with the strongest models on few-shot view synthesis, and can generate coherent 30 second videos from a single image without any explicit geometric information.

Reading the bibliography…