Fetching the paper…

Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding · Around