2023

MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria

Ge, Wentao, Chen, Shunian, Chen, Guiming Hardy et al.

Understand

Multimodal large language models (MLLMs) have broadened the scope of AI applications.

  • Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating queries without considering user experiences, inadequately addressing the nuances of creative and associative multimodal tasks.
  • However, the open-ended and subjective nature of such tasks poses a significant challenge to the evaluation methodology, where it is difficult to define the ground-truth answers for them.
  • To this end, in our paper, we propose a new evaluation paradigm for MLLMs, which is evaluating MLLMs with per-sample criteria using potent MLLM as the judge.

Reading the bibliography…