Fetching the paper…
Reading the bibliography…
The rapid development of multimodal large language models (MLLMs) raises the question of how they compare to human performance.
Rouge: A package for automatic evaluation of summaries
Lin C Y · 2004
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin T Y, Maire M, Belongie S, et al · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol S, Agrawal A, Lu J, et al · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal Y, Khot T, Summers-Stay D, et al · 2017
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino K, Rastegari M, Farhadi A, et al · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson D A, Manning C D · 2019
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford A, Kim J W, Hallacy C, et al · 2021
Earlier work this paper cites.
Align before fuse: Vision and language representation learning with momentum distillation
Li J, Selvaraju R, Gotmare A, et al · 2021
Earlier work this paper cites.
Chatgpt: Optimizing language models for dialogue
OpenAI · 2022
Earlier work this paper cites.
GLM: General language model pretraining with autoregressive blank infilling
Du Z, Qian Y, Liu X, et al · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu P, Mishra S, Xia T, et al · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima T, Gu S S, Reid M, et al · 2022
Earlier work this paper cites.
Gpt-4 technical report, 2023
Achiam J, Adler S, Agarwal S, et al · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models, 2023
Team G, Anil R, Borgeaud S, et al · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models, 2023
Touvron H, Lavril T, Izacard G, et al · 2023
Earlier work this paper cites.
Large multilingual models pivot zero-shot multimodal learning across languages
Hu J, Yao Y, Wang C, et al · 2023
Earlier work this paper cites.
Chinese llava
LinkSoul-AI · 2023
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023
Bai J, Bai S, Yang S, et al · 2023
Earlier work this paper cites.
Gpt-4v(ision) system card
OpenAI · 2023
Earlier work this paper cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li J, Li D, Savarese S, et al · 2023
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai W, Li J, Li D, et al · 2023
Earlier work this paper cites.
Zhang P, Wang X D B, Cao Y, et al · 2023
Earlier work this paper cites.
Touchstone: Evaluating vision-language models by language models
Bai S, Yang S, Bai J, et al · 2023
Cited alongside, same era.
SEED-Bench-2: Benchmarking multimodal large language models
Li B, Ge Y, Ge Y, et al · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Li Y, Du Y, Zhou K, et al · 2023
Cited alongside, same era.
M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models
Zhang W, Aljunied M, Gao C, et al · 2023
Cited alongside, same era.
Scigraphqa: A large-scale synthetic multi-turn question-answering dataset for scientific graphs
Li S, Tajbakhsh N · 2023
Cited alongside, same era.
Scieval: A multi-level large language model evaluation benchmark for scientific research
Sun L, Han Y, Zhao Z, et al · 2024
Closest in time.
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Huang Y, Bai Y, Zhu Z, et al · 2024
Closest in time.
Visual instruction tuning
Liu H, Li C, Wu Q, et al · 2024
Closest in time.
MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Zhu D, Chen J, Shen X, et al · 2024
Closest in time.
MM-vet: Evaluating large multimodal models for integrated capabilities
Yu W, Yang Z, Li L, et al · 2024
Closest in time.
Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Guan T, Liu F, Wu X, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MOSS: An open conversational large language model
Sun T, Zhang X, He Z, et al · 2024
Cited alongside, same era.
Qwen2 technical report, 2024
Team Q · 2024
Cited alongside, same era.
Internlm2 technical report, 2024
Cai Z, Cao M, Chen H, et al · 2024
Cited alongside, same era.
Depression diagnosis dialogue simulation: Self-improving psychiatrist with tertiary memory
Lan K, Jin B, Zhu Z, et al · 2024
Cited alongside, same era.
Ibsen: Director-actor agent collaboration for controllable and interactive drama script generation
Han S, Chen L, Lin L M, et al · 2024
Cited alongside, same era.
Rejection improves reliability: Training llms to refuse unknown questions using rl from knowledge feedback
Xu H, Zhu Z, Zhang S, et al · 2024
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Chen Z, Wu J, Wang W, et al · 2024
Cited alongside, same era.
Closest in time.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu P, Bansal H, Xia T, et al · 2024
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models
Zhong W, Cui R, Guo Y, et al · 2024
Closest in time.
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue X, Ni Y, Zhang K, et al · 2024
Closest in time.
Scibench: Evaluating college-level scientific problem-solving abilities of large language models
Wang X, Hu Z, Lu P, et al · 2024
Closest in time.
CMMU: A benchmark for chinese multi-modal multi-type question understanding and reasoning
He Z, Wu X, Zhou P, et al · 2024
Closest in time.
GAOKAO-MM: A Chinese human-level benchmark for multimodal models evaluation
Zong Y, Qiu X · 2024
Closest in time.
AlignMMBench: Evaluating chinese multimodal alignment in large vision-language models
Wu Y, Yu W, Cheng Y, et al · 2024
Closest in time.
CMMMU: A chinese massive multi-discipline multimodal understanding benchmark
Zhang G, Du X, Chen B, et al · 2024
Closest in time.
DFM: Dialogue foundation model for universal large-scale dialogue-oriented task learning
Chen Z, Ma D, Li H, et al · 2025
Closest in time.
Developing chemdfm as a large language foundation model for chemistry
Zhao Z, Ma D, Chen L, et al · 2025
Closest in time.
Reducing tool hallucination via reliability alignment
Xu H, Zhu Z, Pan L, et al · 2025
Closest in time.
Reasoning-driven retrosynthesis prediction with large language models via reinforcement learning
Zhang S, Li H, Chen L, et al · 2025
Closest in time.
MobA: Multifaceted memory-enhanced adaptive planning for efficient mobile task automation
Zhu Z, Tang H, Li Y, et al · 2025
Closest in time.
Mmbench: Is your multi-modal model an all-around player?
Liu Y, Duan H, Zhang Y, et al · 2025
Closest in time.
MLLM-bench: Evaluating multimodal llms with per-sample criteria
Ge W, Chen S, Chen H, et al · 2025
Closest in time.