2024

TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation

Feng, Weixi, Li, Jiachen, Saxon, Michael et al.

Understand

Video generation has many unique challenges beyond those of image generation.

  • The temporal dimension introduces extensive possible variations across frames, over which consistency and continuity may be violated.
  • In this study, we move beyond evaluating simple actions and argue that generated videos should incorporate the emergence of new concepts and their relation transitions like in real-world videos as time progresses.
  • To assess the Temporal Compositionality of video generation models, we propose TC-Bench, a benchmark of meticulously crafted text prompts, corresponding ground truth videos, and robust evaluation metrics.

Reading the bibliography…