2024

xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Ryoo, Michael S., Zhou, Honglu, Kendre, Shrikant et al.

Understand

We present xGen-MM-Vid (BLIP-3-Video): a multimodal language model for videos, particularly designed to efficiently capture temporal information over multiple frames.

  • BLIP-3-Video takes advantage of the 'temporal encoder' in addition to the conventional visual tokenizer, which maps a sequence of tokens over multiple frames into a compact set of visual tokens.
  • This enables BLIP3-Video to use much fewer visual tokens than its competing models (e.g., 32 vs.
  • 4608 tokens).

Reading the bibliography…