2024

DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes

Song, Zhende, Wang, Chenchen, Sheng, Jiamu et al.

Understand

Recent large vision-language models (LVLMs) for video understanding are primarily fine-tuned with various videos scraped from online platforms.

  • Existing datasets, such as ActivityNet, require considerable human labor for structuring and annotation before effectively utilized for tuning LVLMs.
  • While current LVLMs are primarily trained on existing datasets in broad, general-purpose settings, adapting them to specific downstream scenarios remains challenging, as collecting and annotating task-specific videos is highly labor-intensive and time-consuming.
  • To address this issue, we propose a three-stage framework named DreamFrame for automatically generating style-consistent keyframes and corresponding question-answer (QA) pairs to support LVLM instruction tuning.

Reading the bibliography…