Fetching the paper…

OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model · Around