2022

Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

Qian, Rui, Li, Yeqing, Xu, Zheng et al.

Understand

Utilizing vision and language models (VLMs) pre-trained on large-scale image-text pairs is becoming a promising paradigm for open-vocabulary visual recognition.

  • In this work, we extend this paradigm by leveraging motion and audio that naturally exist in video.
  • We present \textbf{MOV}, a simple yet effective method for \textbf{M}ultimodal \textbf{O}pen-\textbf{V}ocabulary video classification.
  • In MOV, we directly use the vision encoder from pre-trained VLMs with minimal modifications to encode video, optical flow and audio spectrogram.

Reading the bibliography…