Fetching the paper…
Reading the bibliography…
The issue of hallucinations is a prevalent concern in existing Large Vision-Language Models (LVLMs).
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., and Farhadi, A · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al · 2017
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P · 2019
Earlier work this paper cites.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., et al · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Rstnet: Captioning with adaptive attention on visual and non-visual words
Zhang, X., Sun, X., Luo, Y., Ji, J., Zhou, Y., Wu, Y., Huang, F., and Ji, R · 2021
Earlier work this paper cites.
Glm: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J · 2022
Earlier work this paper cites.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Masry, A., Long, D., Tan, J. Q., Joty, S., and Hoque, E · 2022
Earlier work this paper cites.
Difnet: Boosting visual information flow for image captioning
Wu, M., Zhang, X., Sun, X., Zhou, Y., Chen, C., Gu, J., Sun, X., and Ji, R · 2022
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Knvqa: A benchmark for evaluation knowledge-based vqa
Cheng, S., Zhang, S., Wu, J., and Lan, M · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning. arxiv 2023
Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Earlier work this paper cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., Wu, Y., and Ji, R · 2023
Cited alongside, same era.
Llama-adapter v2: Parameter-efficient visual instruction model
Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., Li, H., and Qiao, Y · 2023
Cited alongside, same era.
Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models, 2023
Guan, T., Liu, F., Wu, X., Xian, R., Li, Z., Liu, X., Wang, X., Chen, L., Huang, F., Yacoob, Y., Manocha, D., and Zhou, T · 2023
Cited alongside, same era.
Detecting and preventing hallucinations in large vision language models
Gunjal, A., Yin, J., and Bas, E · 2023
Cited alongside, same era.
Internlm: A multilingual language model with progressively enhanced capabilities
Team, I · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Behind the magic, merlim: Multi-modal evaluation benchmark for large image-language models
Villa, A., Alcázar, J. C. L., Soto, A., and Ghanem, B · 2023
Later among the works it cites.
Woodpecker: Hallucination correction for multimodal large language models
Yin, S., Fu, C., Zhao, S., Xu, T., Wang, H., Sui, D., Shen, Y., Li, K., Sun, X., and Chen, E · 2023
Later among the works it cites.
Mm-vet: Evaluating large multimodal models for integrated capabilities
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hu, H., Zhang, J., Zhao, M., and Sun, Z · 2023
Cited alongside, same era.
Faithscore: Evaluating hallucinations in large vision-language models
Jing, L., Li, R., Chen, Y., Jia, M., and Du, X · 2023
Cited alongside, same era.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al · 2023
Cited alongside, same era.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Leng, S., Zhang, H., Chen, G., Li, X., Lu, S., Miao, C., and Bing, L · 2023
Cited alongside, same era.
FACTUAL: A benchmark for faithful and consistent textual scene graph parsing
Li, Z., Chai, Y., Zhuo, T. Y., Qu, L., Haffari, G., Li, F., Ji, D., and Tran, Q. H · 2023
Cited alongside, same era.
Lu, J., Gan, R., Zhang, D., Wu, X., Wu, Z., Sun, R., Zhang, J., Zhang, P., and Song, Y · 2023
Cited alongside, same era.
Cheap and quick: Efficient vision-language instruction tuning for large language models
Luo, G., Zhou, Y., Ren, T., Chen, S., Sun, X., and Ji, R · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2023
Later among the works it cites.
Halle-switch: Controlling object hallucination in large vision language models
Zhai, B., Yang, S., Xu, C., Shen, S., Keutzer, K., and Li, M · 2023
Later among the works it cites.
Gpt4roi: Instruction tuning large language model on region-of-interest
Zhang, S., Sun, P., Chen, S., Xiao, M., Shao, W., Zhang, W., Chen, K., and Luo, P · 2023
Later among the works it cites.
Beyond hallucinations: Enhancing lvlms through hallucination-aware direct preference optimization
Zhao, Z., Wang, B., Ouyang, L., Dong, X., Wang, J., and He, C · 2023
Later among the works it cites.
Analyzing and mitigating object hallucination in large vision-language models
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H · 2023
Later among the works it cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J · 2024
Closest in time.
Mitigating fine-grained hallucination by fine-tuning large vision-language models with caption rewrites
Wang, L., He, J., Li, S., Liu, N., and Lim, E.-P · 2024
Closest in time.
Toward open-set human object interaction detection
Wu, M., Liu, Y., Ji, J., Sun, X., and Ji, R · 2024
Closest in time.