Fetching the paper…
Reading the bibliography…
Large Vision-Language Models (LVLMs) can reason effectively over both textual and visual inputs, but they tend to hallucinate syntactically coherent yet visually ungrounded contents.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Object hallucination in image captioning
Rohrbach, A., Hendricks, L. A., Burns, K., Darrell, T., and Saenko, K · 2018
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F · 2022
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Li, X. L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T., Zettlemoyer, L., and Lewis, M · 2022
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R · 2023
Earlier work this paper cites.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Earlier work this paper cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., et al · 2023
Earlier work this paper cites.
Evaluating object hallucination in large vision-language models
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, X., and Wen, J.-R · 2023
Earlier work this paper cites.
Video-llava: Learning united visual representation by alignment before projection
Lin, B., Zhu, B., Ye, Y., Ning, M., Jin, P., and Yuan, L · 2023
Earlier work this paper cites.
Mitigating hallucination in large multi-modal models via robust instruction tuning
Liu, F., Lin, K., Li, L., Wang, J., Yacoob, Y., and Wang, L · 2023
Earlier work this paper cites.
Trusting your evidence: Hallucinate less with context-aware decoding
Shi, W., Han, X., Lewis, M., Tsvetkov, Y., Zettlemoyer, L., and Yih, S. W.-t · 2023
Cited alongside, same era.
Aligning large multimodal models with factually augmented rlhf
Sun, Z., Shen, S., Cao, S., Liu, H., Li, C., Shen, Y., Gan, C., Gui, L.-Y., Wang, Y.-X., Yang, Y., et al · 2023
Cited alongside, same era.
Activation addition: Steering language models without optimization, 2023
Turner, A. M., Thiergart, L., Udell, D., Leech, G., Mini, U., and MacDiarmid, M · 2023
Cited alongside, same era.
Label words are anchors: An information flow perspective for understanding in-context learning
Wang, L., Li, L., Dai, D., Chen, D., Zhou, H., Meng, F., Zhou, J., and Sun, X · 2023
Cited alongside, same era.
Woodpecker: Hallucination correction for multimodal large language models
Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation
Huang, Q., Dong, X., Zhang, P., Wang, B., He, C., Wang, J., Lin, D., Zhang, W., and Yu, N · 2024
Later among the works it cites.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Later among the works it cites.
Hallucination augmented contrastive learning for multimodal large language model
Jiang, C., Xu, H., Dong, M., Chen, J., Ye, W., Yan, M., Ye, Q., Zhang, J., Huang, F., and Zhang, S · 2024
Later among the works it cites.
Lisa: Reasoning segmentation via large language model
Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., and Jia, J · 2024
Later among the works it cites.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Leng, S., Zhang, H., Chen, G., Li, X., Lu, S., Miao, C., and Bing, L · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yin, S., Fu, C., Zhao, S., Xu, T., Wang, H., Sui, D., Shen, Y., Li, K., Sun, X., and Chen, E · 2023
Cited alongside, same era.
Analyzing and mitigating object hallucination in large vision-language models
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Cited alongside, same era.
Representation engineering: A top-down approach to ai transparency, 2023
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., Goel, S., Li, N., Byun, M. J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., Kolter, J. Z., and Hendrycks, D · 2023
Cited alongside, same era.
Hallucination of multimodal large language models: A survey
Bai, Z., Wang, P., Xiao, T., He, T., Han, Z., Zhang, Z., and Shou, M. Z · 2024
Cited alongside, same era.
Halc: Object hallucination reduction via adaptive focal-contrast decoding
Chen, Z., Zhao, Z., Luo, H., Yao, H., Li, B., and Zhou, J · 2024
Cited alongside, same era.
Dola: Decoding by contrasting layers improves factuality in large language models
Chuang, Y.-S., Xie, Y., Luo, H., Kim, Y., Glass, J. R., and He, P · 2024
Cited alongside, same era.
Multi-modal hallucination control by visual information grounding
Favero, A., Zancato, L., Trager, M., Choudhary, S., Perera, P., Achille, A., Swaminathan, A., and Soatto, S · 2024
Cited alongside, same era.
Later among the works it cites.
Li, Z., Xu, Z., Han, L., Gao, Y., Wen, S., Liu, D., Wang, H., and Metaxas, D. N · 2024
Later among the works it cites.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024b
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Later among the works it cites.
Eagle: Exploring the design space for multimodal llms with mixture of encoders
Shi, M., Liu, F., Wang, S., Liao, S., Radhakrishnan, S., Huang, D.-A., Yin, H., Sapra, K., Yacoob, Y., Shi, H., et al · 2024
Later among the works it cites.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., and Xie, S · 2024
Later among the works it cites.
Don’t miss the forest for the trees: Attentional vision calibration for large vision language models
Woo, S., Kim, D., Jang, J., Choi, Y., and Kim, C · 2024
Later among the works it cites.
Gpt4tools: Teaching large language model to use tools via self-instruction
Yang, R., Song, L., Li, Y., Zhao, S., Ge, Y., Li, X., and Shan, Y · 2024
Later among the works it cites.
Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data
Yu, Q., Li, J., Wei, L., Pang, L., Ye, W., Qin, B., Tang, S., Tian, Q., and Zhuang, Y · 2024
Later among the works it cites.
Less is more: Mitigating multimodal hallucination from an eos decision perspective
Yue, Z., Zhang, L., and Jin, Q · 2024
Later among the works it cites.