2024

LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Chen, Yukang, Xue, Fuzhao, Li, Dacheng et al.

Understand

Long-context capability is critical for multi-modal foundation models, especially for long video understanding.

  • We introduce LongVILA, a full-stack solution for long-context visual-language models by co-designing the algorithm and system.
  • For model training, we upgrade existing VLMs to support long video understanding by incorporating two additional stages, i.e., long context extension and long video supervised fine-tuning.
  • However, training on long video is computationally and memory intensive.

Reading the bibliography…