Fetching the paper…

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations · Around