Fetching the paper…

ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts · Around