Fetching the paper…
Reading the bibliography…
While large vision-language models (LVLMs) have demonstrated impressive capabilities in interpreting multi-modal contexts, they invariably suffer from object hallucinations (OH).
Longman grammar of spoken and written english, 2000
Biber, D., Johansson, S., Leech, G., Conrad, S., and Finegan, E · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Guided open vocabulary image captioning with constrained beam search
Anderson, P., Fernando, B., Johnson, M., and Gould, S · 2016
Earlier work this paper cites.
spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing
Honnibal, M. and Montani, I · 2017
Earlier work this paper cites.
Object hallucination in image captioning
Rohrbach, A., Hendricks, L. A., Burns, K., Darrell, T., and Saenko, K · 2018
Earlier work this paper cites.
Borgeaud, S. and Emerson, G · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2020
Earlier work this paper cites.
On the limitations of multimodal vaes
Daunhawer, I., Sutter, T. M., Chin-Cheong, K., Palumbo, E., and Vogt, J. E · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Let there be a clock on the beach: Reducing object hallucination in image captioning
Biten, A. F., Gómez, L., and Karatzas, D · 2022
Cited alongside, same era.
Plausible may not be faithful: Probing object hallucination in vision-language pre-training
Dai, W., Liu, Z., Ji, Z., Su, D., and Fung, P · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Minigpt-v2: large language model as a unified interface for vision-language multi-task learning
Chen, J., Zhu, D., Shen, X., Li, X., Liu, Z., Zhang, P., Krishnamoorthi, R., Chandra, V., Xiong, Y., and Elhoseiny, M · 2023
Cited alongside, same era.
Dola: Decoding by contrasting layers improves factuality in large language models
Evaluating object hallucination in large vision-language models
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J.-R · 2023
Later among the works it cites.
Scaling open-vocabulary object detection
Minderer, M., Gritsenko, A., and Houlsby, N · 2023
Later among the works it cites.
Trusting your evidence: Hallucinate less with context-aware decoding
Shi, W., Han, X., Lewis, M., Tsvetkov, Y., Zettlemoyer, L., and Yih, S. W.-t · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
How many unicorns are in this image? a safety evaluation benchmark for vision llms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chuang, Y.-S., Xie, Y., Luo, H., Kim, Y., Glass, J., and He, P · 2023
Cited alongside, same era.
Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges
Cui, C., Zhou, Y., Yang, X., Wu, S., Zhang, L., Zou, J., and Yao, H · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B. A., Fung, P., and Hoi, S. C. H · 2023
Cited alongside, same era.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., et al · 2023
Cited alongside, same era.
The benefits of bad advice: Autocontrastive decoding across model layers
Gera, A., Friedman, R., Arviv, O., Gunasekara, C., Sznajder, B., Slonim, N., and Shnarch, E · 2023
Cited alongside, same era.
Hallusionbench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models
Guan, T., Liu, F., Li, X. W. R. X. Z., Wang, X. L. X., Yacoob, L. C. F. H. Y., and Zhou, D. M. T · 2023
Cited alongside, same era.
Detecting and preventing hallucinations in large vision language models
Gunjal, A., Yin, J., and Bas, E · 2023
Cited alongside, same era.
Huang, Q., Dong, X., Zhang, P., Wang, B., He, C., Wang, J., Lin, D., Zhang, W., and Yu, N · 2023
Cited alongside, same era.
Tu, H., Cui, C., Wang, Z., Zhou, Y., Zhao, B., Han, J., Zhou, W., Yao, H., and Xie, C · 2023
Later among the works it cites.
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Ye, Q., Xu, H., Ye, J., Yan, M., Liu, H., Qian, Q., Zhang, J., Huang, F., and Zhou, J · 2023
Later among the works it cites.
Woodpecker: Hallucination correction for multimodal large language models
Yin, S., Fu, C., Zhao, S., Xu, T., Wang, H., Sui, D., Shen, Y., Li, K., Sun, X., and Chen, E · 2023
Later among the works it cites.
Halle-switch: Controlling object hallucination in large vision language models
Zhai, B., Yang, S., Xu, C., Shen, S., Keutzer, K., and Li, M · 2023
Later among the works it cites.
Analyzing and mitigating object hallucination in large vision-language models
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Scaling open-vocabulary object detection
Minderer, M., Gritsenko, A., and Houlsby, N · 2024
Closest in time.
Wang, X., Zhou, Y., Liu, X., Lu, H., Xu, Y., He, F., Yoon, J., Lu, T., Bertasius, G., Bansal, M., et al · 2024
Closest in time.
Rankclip: Ranking-consistent language-image pretraining
Zhang, Y., Zhao, Z., Chen, Z., Feng, Z., Ding, Z., and Sun, Y · 2024
Closest in time.