2022

Clover: Towards A Unified Video-Language Alignment and Fusion Model

Huang, Jingjia, Li, Yinan, Feng, Jiashi et al.

Understand

Building a universal Video-Language model for solving various video understanding tasks (\emph{e.g.}, text-video retrieval, video question answering) is an open challenge to the machine learning field.

  • Towards this goal, most recent works build the model by stacking uni-modal and cross-modal feature encoders and train it with pair-wise contrastive pre-text tasks.
  • Though offering attractive generality, the resulted models have to compromise between efficiency and performance.
  • They mostly adopt different architectures to deal with different downstream tasks.

Reading the bibliography…