2023

Valley: Video Assistant with Large Language model Enhanced abilitY

Luo, Ruipu, Zhao, Ziwang, Yang, Min et al.

Understand

Large Language Models (LLMs), with remarkable conversational capability, have emerged as AI assistants that can handle both visual and textual modalities.

  • However, their effectiveness in joint video and language understanding has not been extensively explored.
  • In the paper, we introduce Valley, a multi-modal foundation model that is designed to enable enhanced video comprehension and instruction-following capabilities.
  • To this end, we construct two datasets, namely Valley-702k and Valley-instruct-73k, to cover a diverse range of video-text alignment and video-based instruction tasks, such as multi-shot captions, long video descriptions, action recognition, causal inference, etc.

Reading the bibliography…