Fetching the paper…

VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation · Around