Fetching the paper…
Reading the bibliography…
Despite their impressive capabilities, multimodal large language models (MLLMs) are prone to hallucinations, i.e., the generated content that is nonsensical or unfaithful to input sources.
Entropy, relative entropy and mutual information
Cover, T. M., Thomas, J. A., et al · 1991
Earlier work this paper cites.
Memory representations in natural tasks
Ballard, D. H., Hayhoe, M. M., and Pelz, J. B · 1995
Earlier work this paper cites.
Visual search has no memory
Horowitz, T. S. and Wolfe, J. M · 1998
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Information dropout: Learning optimal representations through noisy computation
Achille, A. and Soatto, S · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
Object hallucination in image captioning
Rohrbach, A., Hendricks, L. A., Burns, K., Darrell, T., and Saenko, K · 2018
Earlier work this paper cites.
Depth-adaptive transformer
Elbayad, M., Gu, J., Grave, E., and Auli, M · 2020
Earlier work this paper cites.
Kao, W.-T., Wu, T.-H., Chi, P.-H., Hsieh, C.-C., and Lee, H.-Y · 2020
Earlier work this paper cites.
Evolving normalization-activation layers
Liu, H., Brock, A., Simonyan, K., and Le, Q · 2020
Earlier work this paper cites.
On the limitations of multimodal vaes
Daunhawer, I., Sutter, T. M., Chin-Cheong, K., Palumbo, E., and Vogt, J. E · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Earlier work this paper cites.
Confident adaptive language modeling
Schuster, T., Fisch, A., Gupta, J., Dehghani, M., Bahri, D., Tran, V., Tay, Y., and Metzler, D · 2022
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Dola: Decoding by contrasting layers improves factuality in large language models
Chuang, Y.-S., Xie, Y., Luo, H., Kim, Y., Glass, J. R., and He, P · 2023
Earlier work this paper cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., et al · 2023
Cited alongside, same era.
Huang, Q., Dong, X., Zhang, P., Wang, B., He, C., Wang, J., Lin, D., Zhang, W., and Yu, N · 2023
Cited alongside, same era.
Exposing and mitigating spurious correlations for cross-modal retrieval
Kim, J. M., Koepke, A., Schmid, C., and Akata, Z · 2023
Cited alongside, same era.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Leng, S., Zhang, H., Chen, G., Li, X., Lu, S., Miao, C., and Bing, L · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Generating images with multimodal language models
Koh, J. Y., Fried, D., and Salakhutdinov, R. R · 2024
Closest in time.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Leng, S., Zhang, H., Chen, G., Li, X., Lu, S., Miao, C., and Bing, L · 2024
Closest in time.
Llava-next: Stronger llms supercharge multimodal capabilities in the wild
Li, B., Zhang, K., Zhang, H., Guo, D., Zhang, R., Li, F., Zhang, Y., Liu, Z., and Li, C · 2024
Closest in time.
Has multimodal learning delivered universal intelligence in healthcare? a comprehensive survey
Lin, Q., Zhu, Y., Mei, X., Huang, L., Ma, J., He, K., Peng, Z., Cambria, E., and Feng, M · 2024
Closest in time.
Neo, D. and Chen, T · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, X., and Wen, J.-R · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Sight beyond text: Multi-modal training enhances llms in truthfulness and ethics
Tu, H., Zhao, B., Wei, C., and Xie, C · 2023
Cited alongside, same era.
Cogvlm: Visual expert for pretrained language models
Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X., et al · 2023
Cited alongside, same era.
A survey on multimodal large language models
Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., and Chen, E · 2023
Cited alongside, same era.
Analyzing and mitigating object hallucination in large vision-language models
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H · 2023
Cited alongside, same era.
Hallucination of multimodal large language models: A survey
Bai, Z., Wang, P., Xiao, T., He, T., Han, Z., Zhang, Z., and Shou, M. Z · 2024
Cited alongside, same era.
Holistic autonomous driving understanding by bird’s-eye-view injected multi-modal large models
Ding, X., Han, J., Xu, H., Liang, X., Zhang, W., and Li, X · 2024
Cited alongside, same era.
Alleviating hallucination in large vision-language models with active retrieval augmentation
Qu, X., Chen, Q., Wei, W., Sun, J., and Dong, J · 2024
Closest in time.
Trusting your evidence: Hallucinate less with context-aware decoding
Shi, W., Han, X., Lewis, M., Tsvetkov, Y., Zettlemoyer, L., and Yih, W.-t · 2024
Closest in time.
The evolution of multimodal model architectures
Wadekar, S. N., Chaurasia, A., Chadha, A., and Culurciello, E · 2024
Closest in time.
Mitigating hallucinations in large vision-language models with instruction contrastive decoding
Wang, X., Pan, J., Ding, L., and Biemann, C · 2024
Closest in time.
Next-gpt: Any-to-any multimodal llm
Wu, S., Fei, H., Qu, L., Ji, W., and Chua, T.-S · 2024
Closest in time.
Mitigating object hallucination via concentric causal attention
Xing, Y., Li, Y., Laptev, I., and Lu, S · 2024
Closest in time.
Im-rag: Multi-round retrieval-augmented generation through learning inner monologues
Yang, D., Rao, J., Chen, K., Guo, X., Zhang, Y., Yang, J., and Zhang, Y · 2024
Closest in time.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Huang, X., Wang, Z., Sheng, L., Bai, L., et al · 2024
Closest in time.
Seeing clearly by layer two: Enhancing attention heads to alleviate hallucination in lvlms
Zhang, X., Quan, Y., Gu, C., Shen, C., Yuan, X., Yan, S., Cheng, H., Wu, K., and Ye, J · 2024
Closest in time.
Zheng, K., Chen, J., Yan, Y., Zou, X., and Hu, X · 2024
Closest in time.
Analyzing and mitigating object hallucination in large vision-language models
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H · 2024
Closest in time.
Realrag: Retrieval-augmented realistic image generation via self-reflective contrastive learning
Lyu, Y., Zheng, X., Jiang, L., Yan, Y., Zou, X., Zhou, H., Zhang, L., and Hu, X · 2025
Closest in time.