Fetching the paper…
Reading the bibliography…
Despite their remarkable potential, Large Vision-Language Models (LVLMs) still face challenges with object hallucination, a problem where their generated outputs mistakenly incorporate objects that do not actually exist.
Nltk: the natural language toolkit
S. Bird · 2006
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Object hallucination in image captioning
A. Rohrbach, L. A. Hendricks, K. Burns, T. Darrell, and K. Saenko · 2018
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
X. L. Li, A. Holtzman, D. Fried, P. Liang, J. Eisner, T. Hashimoto, L. Zettlemoyer, and M. Lewis · 2022
Earlier work this paper cites.
Shikra: Unleashing multimodal llm’s referential dialogue magic
K. Chen, Z. Zhang, W. Zeng, R. Zhang, F. Zhu, and R. Zhao · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, et al · 2023
Earlier work this paper cites.
Volcano: mitigating multimodal hallucination through self-feedback guided revision
S. Lee, S. H. Park, Y. Jo, and M. Seo · 2023
Earlier work this paper cites.
Evaluating object hallucination in large vision-language models
Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen · 2023
Earlier work this paper cites.
Aligning large multimodal models with factually augmented rlhf
Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y. Shen, C. Gan, L.-Y. Gui, Y.-X. Wang, Y. Yang, et al · 2023
Earlier work this paper cites.
Efficient streaming language models with attention sinks
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis · 2023
Cited alongside, same era.
Analyzing and mitigating object hallucination in large vision-language models
Y. Zhou, C. Cui, J. Yoon, L. Zhang, Z. Deng, C. Finn, M. Bansal, and H. Yao · 2023
Cited alongside, same era.
Hallucination of multimodal large language models: A survey
Z. Bai, P. Wang, T. Xiao, T. He, Z. Han, Z. Zhang, and M. Z. Shou · 2024
Cited alongside, same era.
Dola: Decoding by contrasting layers improves factuality in large language models
Y.-S. Chuang, Y. Xie, H. Luo, Y. Kim, J. R. Glass, and P. He · 2024
Cited alongside, same era.
Multi-modal hallucination control by visual information grounding
A. Favero, L. Zancato, M. Trager, S. Choudhary, P. Perera, A. Achille, A. Swaminathan, and S. Soatto · 2024
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024b
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Later among the works it cites.
Towards interpreting visual information processing in vision-language models
C. Neo, L. Ong, P. Torr, M. Geva, D. Krueger, and F. Barez · 2024
Later among the works it cites.
Interpreting gpt: The logit lens, August 2020
nostalgebraist · 2024
Later among the works it cites.
OpenAI and A. et al · 2024
Later among the works it cites.
Dopra: Decoding over-accumulation penalization and re-allocation in specific weighting layer
J. Wei and X. Zhang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Damro: Dive into the attention mechanism of lvlm to reduce object hallucination
X. Gong, T. Ming, X. Wang, and Z. Wei · 2024
Cited alongside, same era.
Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob, et al · 2024
Cited alongside, same era.
Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation
Q. Huang, X. Dong, P. Zhang, B. Wang, C. He, J. Wang, D. Lin, W. Zhang, and N. Yu · 2024
Cited alongside, same era.
Interpreting and editing vision-language representations to mitigate hallucinations
N. Jiang, A. Kachinthaya, S. Petryk, and Y. Gandelsman · 2024
Cited alongside, same era.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
S. Leng, H. Zhang, G. Chen, X. Li, S. Lu, C. Miao, and L. Bing · 2024
Cited alongside, same era.
Improved baselines with visual instruction tuning
H. Liu, C. Li, Y. Li, and Y. J. Lee
Cited in the paper.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee
Cited in the paper.
S. Woo, D. Kim, J. Jang, Y. Choi, and C. Kim · 2024
Later among the works it cites.
Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data
Q. Yu, J. Li, L. Wei, L. Pang, W. Ye, B. Qin, S. Tang, Q. Tian, and Y. Zhuang · 2024
Later among the works it cites.
Self-introspective decoding: Alleviating hallucinations for large vision-language models
F. Huo, W. Xu, Z. Zhang, H. Wang, Z. Chen, and P. Zhao · 2025
Closest in time.
Paying more attention to image: A training-free method for alleviating hallucination in lvlms
S. Liu, K. Zheng, and W. Chen · 2025
Closest in time.