Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) mimic human perception and reasoning system by integrating powerful Large Language Models (LLMs) with various modality encoders (e.g., vision, audio), positioning LLMs as the "brain" and various modality encoders as sensory organs.
Microsoft coco: Common objects in context
T.-Y. Lin et al · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young et al · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam et al · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen et al · 2015
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning, 2016
J. Johnson et al · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani et al · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
C. R. Qi et al · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2017
Y. Goyal et al · 2017
Earlier work this paper cites.
Fvqa: Fact-based visual question answering
P. Wang et al · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
A. Dai et al · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari et al · 2018
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
J. J. Lau et al · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering, 2019
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Towards vqa models that can read
A. Singh et al · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
K. Marino et al · 2019
Earlier work this paper cites.
Raven: A dataset for relational and analogical visual reasoning
C. Zhang et al · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers, 2019
H. Tan and M. Bansal · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
C. D. Kim et al · 2019
Earlier work this paper cites.
Textcaps: a dataset for image captioning with reading comprehension
O. Sidorov et al · 2020
Earlier work this paper cites.
Document visual question answering challenge 2020
M. Mathew et al · 2020
Earlier work this paper cites.
Point-bert: Pre-training 3d point cloud transformers with masked point modeling
X. Yu et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford et al · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann et al · 2021
Earlier work this paper cites.
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
P. Lu et al · 2021
Earlier work this paper cites.
Vilt: Vision-and-language transformer without convolution or region supervision, 2021
W. Kim et al · 2021
Earlier work this paper cites.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Y. Zeng et al · 2021
Earlier work this paper cites.
A-okvqa: A benchmark for visual question answering using world knowledge
D. Schwenk et al · 2022
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
A. Masry et al · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
P. Lu et al · 2022
Earlier work this paper cites.
Scanqa: 3d question answering for spatial scene understanding
D. Azuma et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei et al · 2022
Earlier work this paper cites.
Seed-bench: Benchmarking multimodal llms with generative comprehension
B. Li et al · 2023
Earlier work this paper cites.
A survey on large language models: Applications, challenges, limitations, and practical usage
M. U. Hadi et al · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
H. Touvron et al · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
L. Zheng et al · 2023
Earlier work this paper cites.
Qlora: Efficient finetuning of quantized llms
T. Dettmers et al · 2023
Earlier work this paper cites.
Visual spatial reasoning
F. Liu et al · 2023
Earlier work this paper cites.
Bubogpt: Enabling visual grounding in multi-modal llms
Y. Zhao et al · 2023
Earlier work this paper cites.
Valley: Video assistant with large language model enhanced ability
R. Luo et al · 2023
Earlier work this paper cites.
Robust speech recognition via large-scale weak supervision
A. Radford et al · 2023
Earlier work this paper cites.
Improved baselines with visual instruction tuning
H. Liu et al · 2023
Earlier work this paper cites.
Aligning large multimodal models with factually augmented rlhf
Z. Sun et al · 2023
Earlier work this paper cites.
Learning the user’s deeper preferences for multi-modal recommendation systems
F. Lei et al · 2023
Earlier work this paper cites.
Mmbench: Is your multi-modal model an all-around player?
Y. Liu et al · 2023
Earlier work this paper cites.
Mm-vet: Evaluating large multimodal models for integrated capabilities
W. Yu et al · 2023
Earlier work this paper cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models
C. Fu et al · 2023
Earlier work this paper cites.
H. Liu et al · 2023
Earlier work this paper cites.
Equivariant similarity for vision-language foundation models
T. Wang et al · 2023
Earlier work this paper cites.
Touchstone: Evaluating vision-language models by language models
S. Bai et al · 2023
Earlier work this paper cites.
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Q. Ye et al · 2023
Earlier work this paper cites.
Vl-checklist: Evaluating pre-trained vision-language models with objects, attributes and relations, 2023
T. Zhao et al · 2023
Earlier work this paper cites.
When and why vision-language models behave like bags-of-words, and what to do about it?
M. Yuksekgonul et al · 2023
Earlier work this paper cites.
Q-bench: A benchmark for general-purpose foundation models on low-level vision
H. Wu et al · 2023
Earlier work this paper cites.
On the hidden mystery of ocr in large multimodal models
Y. Liu et al · 2023
Earlier work this paper cites.
Hierarchical multimodal transformers for multipage docvqa
R. Tito et al · 2023
Earlier work this paper cites.
Contextual object detection with multimodal large language models
Y. Zang et al · 2023
Earlier work this paper cites.
What’s” up” with vision-language models? investigating their struggle with spatial reasoning
A. Kamath et al · 2023
Earlier work this paper cites.
Chartbench: A benchmark for complex visual reasoning in charts
Z. Xu et al · 2023
Earlier work this paper cites.
Scigraphqa: A large-scale synthetic multi-turn question-answering dataset for scientific graphs
S. Li and N. Tajbakhsh · 2023
Earlier work this paper cites.
Mmc: Advancing multimodal chart understanding with large-scale instruction tuning
F. Liu et al · 2023
Earlier work this paper cites.
Z. Shi et al · 2023
Earlier work this paper cites.
Benchlmm: Benchmarking cross-style visual capability of large multimodal models
R. Cai et al · 2023
Earlier work this paper cites.
Evaluating object hallucination in large vision-language models
Y. Li et al · 2023
Cited alongside, same era.
Mitigating hallucination in large multi-modal models via robust instruction tuning
F. Liu et al · 2023
Cited alongside, same era.
Evaluation and analysis of hallucination in large vision-language models
J. Wang et al · 2023
Cited alongside, same era.
Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges
C. Cui et al · 2023
Cited alongside, same era.
An llm-free multi-dimensional benchmark for mllms hallucination evaluation
Vlkeb: A large vision-language model knowledge editing benchmark, 2024
H. Huang et al · 2024
Closest in time.
Mc-mke: A fine-grained multimodal knowledge editing benchmark emphasizing modality consistency
J. Zhang et al · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue et al · 2024
Closest in time.
X. Wang et al · 2024
Closest in time.
Marvel: Multidimensional abstraction and reasoning through visual evaluation and learning
Y. Jiang et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Wang et al · 2023
Cited alongside, same era.
Mm-safetybench: A benchmark for safety evaluation of multimodal large language models
X. Liu et al · 2023
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
P. Lu et al · 2023
Cited alongside, same era.
M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models
W. Zhang et al · 2023
Cited alongside, same era.
L. Chen et al · 2023
Cited alongside, same era.
M3dbench: Let’s instruct large models with multi-modal 3d prompts
M. Li et al · 2023
Cited alongside, same era.
Cross-city matters: A multimodal remote sensing benchmark dataset for cross-city semantic segmentation using high-resolution domain adaptation networks
D. Hong et al · 2023
Cited alongside, same era.
Rsgpt: A remote sensing vision language model and benchmark
Y. Hu et al · 2023
Cited alongside, same era.
What is the visual cognition gap between humans and multimodal llms?
X. Cao et al · 2024
Closest in time.
Needle in a multimodal haystack
W. Wang et al · 2024
Closest in time.
Chartx & chartvlm: A versatile benchmark and foundation model for complicated chart reasoning
R. Xia et al · 2024
Closest in time.
Charxiv: Charting gaps in realistic chart understanding in multimodal llms
Z. Wang et al · 2024
Closest in time.
How easy is it to fool your multimodal llms? an empirical analysis on deceptive prompts
Y. Qian et al · 2024
Closest in time.
Y. Liu et al · 2024
Closest in time.
Mm-spubench: Towards better understanding of spurious biases in multimodal llms
W. Ye et al · 2024
Closest in time.
Benchmarking trustworthiness of multimodal large language models: A comprehensive study
Y. Zhang et al · 2024
Closest in time.
Unified hallucination detection for multimodal large language models
X. Chen et al · 2024
Closest in time.
Videohallucer: Evaluating intrinsic and extrinsic hallucinations in large video-language models
Y. Wang et al · 2024
Closest in time.
Visually dehallucinative instruction generation
S. Cha et al · 2024
Closest in time.
Detecting and preventing hallucinations in large vision language models
A. Gunjal et al · 2024
Closest in time.
Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
T. Guan et al · 2024
Closest in time.
Y. Wang et al · 2024
Closest in time.
Visual hallucinations of multi-modal large language models
W. Huang et al · 2024
Closest in time.
The instinctive bias: Spurious images lead to hallucination in mllms
T. Han et al · 2024
Closest in time.
Red teaming visual language models
M. Li et al · 2024
Closest in time.
Single image unlearning: Efficient machine unlearning in multimodal large language models
J. Li et al · 2024
Closest in time.
W. Luo et al · 2024
Closest in time.
Y. Shi et al · 2024
Closest in time.
Cvqa: Culturally-diverse multilingual visual question answering benchmark
D. Romero et al · 2024
Closest in time.
Mm-soc: Benchmarking multimodal large language models in social media platforms
Y. Jin et al · 2024
Closest in time.
Transportationgames: Benchmarking transportation knowledge of (multimodal) large language models
X. Zhang et al · 2024
Closest in time.
Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?
R. Zhang et al · 2024
Closest in time.
Nphardeval4v: A dynamic reasoning benchmark of multimodal large language models
L. Fan et al · 2024
Closest in time.
Scemqa: A scientific college entrance level multimodal question answering benchmark
Z. Liang et al · 2024
Closest in time.
Multi: Multimodal understanding leaderboard with text and images
Z. Zhu et al · 2024
Closest in time.
Measuring multimodal mathematical reasoning with math-vision dataset, 2024
K. Wang et al · 2024
Closest in time.
Is your model really a good math reasoner? evaluating mathematical reasoning with checklist
Z. Zhou et al · 2024
Closest in time.
Cmmmu: A chinese massive multi-discipline multimodal understanding benchmark
G. Zhang et al · 2024
Closest in time.
Peacock: A family of arabic multimodal large language models and benchmarks
F. Alwajih et al · 2024
Closest in time.
Lavy: Vietnamese multimodal large language model
C. Tran and H. L. Thanh · 2024
Closest in time.
Mtvqa: Benchmarking multilingual text-centric visual question answering
J. Tang et al · 2024
Closest in time.
Scifibench: Benchmarking large multimodal models for scientific figure interpretation
J. Roberts et al · 2024
Closest in time.
A. C. Doris et al · 2024
Closest in time.
Asclepius: A spectrum evaluation benchmark for medical multi-modal large language models
W. Wang et al · 2024
Closest in time.
M3d: Advancing 3d medical image analysis with multi-modal large language models
F. Bai et al · 2024
Closest in time.
Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai
P. Chen et al · 2024
Closest in time.
Mobile-agent: Autonomous multi-modal mobile device agent with visual perception
J. Wang et al · 2024
Closest in time.
Visualagentbench: Towards large multimodal models as visual foundation agents
X. Liu et al · 2024
Closest in time.
Egoplan-bench: Benchmarking multimodal large language models for human-level planning, 2024
Y. Chen et al · 2024
Closest in time.
Openeqa: Embodied question answering in the era of foundation models
A. Majumdar et al · 2024
Closest in time.
Ferret-ui: Grounded mobile ui understanding with multimodal llms
K. You et al · 2024
Closest in time.
Crab: Cross-environment agent benchmark for multimodal language model agents
T. Xu et al · 2024
Closest in time.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Z. Yin et al · 2024
Closest in time.
Mmbench-video: A long-form multi-shot benchmark for holistic video understanding
X. Fang et al · 2024
Closest in time.
Sok-bench: A situated video reasoning benchmark with aligned open-world knowledge
A. Wang et al · 2024
Closest in time.
Mvbench: A comprehensive multi-modal video understanding benchmark
K. Li et al · 2024
Closest in time.
Air-bench: Benchmarking large audio-language models via generative comprehension
Q. Yang et al · 2024
Closest in time.
Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech
C.-y. Huang et al · 2024
Closest in time.
Muchomusic: Evaluating music understanding in multimodal audio-language models
B. Weck et al · 2024
Closest in time.
X. Dong et al · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI et al · 2024
Closest in time.
Osprey: Pixel understanding with visual instruction tuning, 2024
Y. Yuan et al · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
H. Liu et al · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
G. Team et al · 2024
Closest in time.
Obelics: An open web-scale filtered dataset of interleaved image-text documents
H. Laurençon et al · 2024
Closest in time.