Fetching the paper…
Reading the bibliography…
Large vision-language models (LVMs) extend large language models (LLMs) with visual perception capabilities, enabling them to process and interpret visual information.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Chen, P.-Y., Zhang, H., Sharma, Y., Yi, J., and Hsieh, C.-J · 2017
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Nesterov, Y. and Spokoiny, V · 2017
Earlier work this paper cites.
Object hallucination in image captioning
Rohrbach, A., Hendricks, L. A., Burns, K., Darrell, T., and Saenko, K · 2018
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Let there be a clock on the beach: Reducing object hallucination in image captioning
Biten, A. F., Gómez, L., and Karatzas, D · 2022
Earlier work this paper cites.
In NeurIPS , 2022
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S., Crowson, K., Schmidt, L., Kaczmarczyk, R., and Jitsev, J · 2022
Earlier work this paper cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
Thrush, T., Jiang, R., Bartolo, M., Singh, A., Williams, A., Kiela, D., and Ross, C · 2022
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., LI, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P. N., and Hoi, S · 2023
Earlier work this paper cites.
Evaluating object hallucination in large vision-language models
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, X., and Wen, J.-R · 2023
Earlier work this paper cites.
Cheap and quick: Efficient vision-language instruction tuning for large language models
Luo, G., Zhou, Y., Ren, T., Chen, S., Sun, X., and Ji, R · 2023
Earlier work this paper cites.
What does clip know about a red circle? visual prompt engineering for vlms
Shtedritski, A., Rupprecht, C., and Vedaldi, A · 2023
Earlier work this paper cites.
On evaluating adversarial robustness of large vision-language models
Zhao, Y., Pang, T., Du, C., Yang, X., LI, C., Cheung, N.-M. M., and Lin, M · 2023
Cited alongside, same era.
Hallucination of multimodal large language models: A survey
Bai, Z., Wang, P., Xiao, T., He, T., Han, Z., Zhang, Z., and Shou, M. Z · 2024
Cited alongside, same era.
On the robustness of large multimodal models against image adversarial attacks
Cui, X., Aparcedo, A., Jang, Y. K., and Lim, S.-N · 2024
Cited alongside, same era.
Language modeling is compression
Deletang, G., Ruoss, A., Duquenne, P.-A., Catt, E., Genewein, T., Mattern, C., Grau-Moya, J., Wenliang, L. K., Aitchison, M., Orseau, L., et al · 2024
Cited alongside, same era.
Multi-modal hallucination control by visual information grounding
Favero, A., Zancato, L., Trager, M., Choudhary, S., Perera, P., Achille, A., Swaminathan, A., and Soatto, S · 2024
Cited alongside, same era.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Leng, S., Zhang, H., Chen, G., Li, X., Lu, S., Miao, C., and Bing, L · 2024
Later among the works it cites.
Llava-onevision: Easy visual task transfer
Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Zhang, P., Li, Y., Liu, Z., et al · 2024
Later among the works it cites.
Ovis: Structural embedding alignment for multimodal large language model
Lu, S., Li, Y., Chen, Q.-G., Xu, Z., Luo, W., Zhang, K., and Ye, H.-J · 2024
Later among the works it cites.
Task bias in contrastive vision-language models
Menon, S., Chandratreya, I. P., and Vondrick, C · 2024
Later among the works it cites.
Can i trust your answer? visually grounded video question answering
Xiao, J., Yao, A., Li, Y., and Chua, T.-S · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Guan, T., Liu, F., Wu, X., Xian, R., Li, Z., Liu, X., Wang, X., Chen, L., Huang, F., Yacoob, Y., Manocha, D., and Zhou, T · 2024
Cited alongside, same era.
Detecting and preventing hallucinations in large vision language models
Gunjal, A., Yin, J., and Bas, E · 2024
Cited alongside, same era.
Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation
Huang, Q., Dong, X., Zhang, P., Wang, B., He, C., Wang, J., Lin, D., Zhang, W., and Yu, N · 2024
Cited alongside, same era.
Exploiting semantic reconstruction to mitigate hallucinations in vision-language models
Kim, M., Kim, M., Bae, J., Choi, S., Kim, S., and Chang, B · 2024
Cited alongside, same era.
Geochat: Grounded large vision-language model for remote sensing
Kuckreja, K., Danish, M. S., Naseer, M., Das, A., Khan, S., and Khan, F. S · 2024
Cited alongside, same era.
What matters when building vision-language models?
Laurençon, H., Tronchon, L., Cord, M., and Sanh, V · 2024
Cited alongside, same era.
Sharegpt4v: Improving large multi-modal models with better captions
Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D
Cited in the paper.
Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models
Xu, P., Shao, W., Zhang, K., Gao, P., Liu, S., Lei, M., Meng, F., Huang, S., Qiao, Y., and Luo, P · 2024
Later among the works it cites.
Beaf: Observing before-after changes to evaluate hallucination in vision-language models
Ye-Bin, M., Hyeon-Woo, N., Choi, W., and Oh, T.-H · 2024
Later among the works it cites.
Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data
Yu, Q., Li, J., Wei, L., Pang, L., Ye, W., Qin, B., Tang, S., Tian, Q., and Zhuang, Y · 2024
Later among the works it cites.
Analyzing and mitigating object hallucination in large vision-language models
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H · 2024
Later among the works it cites.
PerturboLLaVA: Reducing multimodal hallucinations with perturbative visual training
Anonymous · 2025
Closest in time.