Fetching the paper…

Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes · Around