2023

VaQuitA: Enhancing Alignment in LLM-Assisted Video Understanding

Wang, Yizhou, Zhang, Ruiyi, Wang, Haoliang et al.

Understand

Recent advancements in language-model-based video understanding have been progressing at a remarkable pace, spurred by the introduction of Large Language Models (LLMs).

  • However, the focus of prior research has been predominantly on devising a projection layer that maps video features to tokens, an approach that is both rudimentary and inefficient.
  • In our study, we introduce a cutting-edge framework, VaQuitA, designed to refine the synergy between video and textual information.
  • At the data level, instead of sampling frames uniformly, we implement a sampling method guided by CLIP-score rankings, which enables a more aligned selection of frames with the given question.

Reading the bibliography…