Fetching the paper…

VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents · Around