2023

Osprey: Pixel Understanding with Visual Instruction Tuning

Yuan, Yuqian, Li, Wentong, Liu, Jian et al.

Understand

Multimodal large language models (MLLMs) have recently achieved impressive general-purpose vision-language capabilities through visual instruction tuning.

  • However, current MLLMs primarily focus on image-level or box-level understanding, falling short in achieving fine-grained vision-language alignment at pixel level.
  • Besides, the lack of mask-based instruction data limits their advancements.
  • In this paper, we propose Osprey, a mask-text instruction tuning approach, to extend MLLMs by incorporating fine-grained mask regions into language instruction, aiming at achieving pixel-wise visual understanding.

Reading the bibliography…