Fetching the paper…

MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models · Around