Fetching the paper…

LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description · Around