Fetching the paper…
Reading the bibliography…
Large Vision-Language Models (LVLMs) have shown significant capability in vision-language understanding.
2019
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transferable visual models from natural language supervision, in: International conference on machine learning, PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
J. Ho, T. Salimans, Classifier-free diffusion guidance, in: NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021
2021
Earlier work this paper cites.
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, A. Kalyan, Learn to explain: Multimodal reasoning via thought chains for science question answering, Advances in Neural Information Processing Systems 35 (2022) 2507–2521
2022
Earlier work this paper cites.
K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu, et al., A survey on vision transformer, IEEE transactions on pattern analysis and machine intelligence 45 (1) (2022) 87–110
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al., Chain-of-thought prompting elicits reasoning in large language models, Advances in neural information processing systems 35 (2022) 24824–24837
2022
Earlier work this paper cites.
E. Perez, S. Ringer, K. Lukosiute, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kadavath, et al., Discovering language model behaviors with model-written evaluations, in: Findings of the Association for Computational Linguistics: ACL 2023, 2023, pp. 13387–13434
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, J.-R. Wen, Evaluating object hallucination in large vision-language models, in: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 292–305
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. L. Li, A. Holtzman, D. Fried, P. Liang, J. Eisner, T. B. Hashimoto, L. Zettlemoyer, M. Lewis, Contrastive decoding: Open-ended text generation as optimization, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 12286–12312
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, I. Stoica, Efficient memory management for large language model serving with pagedattention, in: Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023
2023
Earlier work this paper cites.
doi:https://doi.org/10.1016/j.neucom.2024.127530
Y. Zhang, C. Zhang, Y. Tang, Z. He, Cross-modal concept learning and inference for vision-language models , Neurocomputing 583 (2024) 127530 · 2024
Cited alongside, same era.
doi:https://doi.org/10.1016/j.neucom.2024.129131
J. Rodriguez-Juan, D. Ortiz-Perez, J. Garcia-Rodriguez, D. Tomás, G. J.Nalepa, Integrating advanced vision-language models for context recognition in risks assessment , Neurocomputing 618 (2025) 129131 · 2024
Cited alongside, same era.
D. Zhang, Y. Yu, J. Dong, C. Li, D. Su, C. Chu, D. Yu, Mm-llms: Recent advances in multimodal large language models, in: Findings of the Association for Computational Linguistics ACL 2024, 2024, pp. 12401–12430
2024
Cited alongside, same era.
C. Cui, Y. Ma, X. Cao, W. Ye, Y. Zhou, K. Liang, J. Chen, J. Lu, Z. Yang, K.-D. Liao, et al., A survey on multimodal large language models for autonomous driving, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 958–979
2024
Cited alongside, same era.
X. Liu, J. Wu, W. Yang, X. Zhou, T. Zhang, Multi-modal attribute prompting for vision-language models, IEEE Transactions on Circuits and Systems for Video Technology (2024)
2024
Closest in time.
M. Wu, J. Ji, O. Huang, J. Li, Y. Wu, X. Sun, R. Ji, Evaluating and analyzing relationship hallucinations in large vision-language models, in: Forty-first International Conference on Machine Learning, 2024
2024
Closest in time.
T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob, et al., Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14375–14385
2024
Closest in time.
Q. Huang, X. Dong, P. Zhang, B. Wang, C. He, J. Wang, D. Lin, W. Zhang, N. Yu, Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13418–13427
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, E. DURMUS, Z. Hatfield-Dodds, S. R. Johnston, S. M. Kravec, et al., Towards understanding sycophancy in language models, in: The Twelfth International Conference on Learning Representations, 2024
2024
Cited alongside, same era.
Y. Qian, H. Zhang, Y. Yang, Z. Gan, How easy is it to fool your multimodal LLMs? an empirical analysis on deceptive prompt , in: Neurips Safe Generative AI Workshop 2024, 2024. URL https://openreview.net/forum?id=BGY6LWN8bh
2024
Cited alongside, same era.
B. Li, Y. Ge, Y. Ge, G. Wang, R. Wang, R. Zhang, Y. Shan, Seed-bench: Benchmarking multimodal large language models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13299–13308
2024
Cited alongside, same era.
W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, L. Wang, Mm-vet: Evaluating large multimodal models for integrated capabilities, in: International conference on machine learning, PMLR, 2024
2024
Cited alongside, same era.
x.ai, Realworldqa: A benchmark for real-world spatial understanding, https://x.ai/blog/grok-1.5v , accessed: 2024-07-09 (2024)
2024
Cited alongside, same era.
W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, S. XiXuan, et al., Cogvlm: Visual expert for pretrained language models, Advances in Neural Information Processing Systems 37 (2024) 121475–121499
2024
Cited alongside, same era.
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al., Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 24185–24198
2024
Cited alongside, same era.
2024
Closest in time.
Z. Chen, Z. Zhao, H. Luo, H. Yao, B. Li, J. Zhou, Halc: Object hallucination reduction via adaptive focal-contrast decoding, in: Forty-first International Conference on Machine Learning, 2024
2024
Closest in time.
S. Leng, H. Zhang, G. Chen, X. Li, S. Lu, C. Miao, L. Bing, Mitigating object hallucinations in large vision-language models through visual contrastive decoding, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13872–13882
2024
Closest in time.
2024
Closest in time.
S. Lee, S. Park, Y. Jo, M. Seo, Volcano: Mitigating multimodal hallucination through self-feedback guided revision, in: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 391–404
2024
Closest in time.
doi:https://doi.org/10.1016/j.neucom.2025.130028
Y. Liu, Y. Deng, A. Liu, Y. Liu, S. Li, Fine-grained multi-modal prompt learning for vision–language models , Neurocomputing 636 (2025) 130028 · 2025
Closest in time.
doi:https://doi.org/10.1016/j.neucom.2025.129457
Z. Wang, M. Li, M. Wu, M.-F. Moens, T. Tuytelaars, Instruction-guided path planning with 3d semantic maps for vision-language navigation , Neurocomputing 625 (2025) 129457 · 2025
Closest in time.
J. Xiao, N. Huang, H. Qin, D. Li, Y. Li, F. Zhu, Z. Tao, J. Yu, L. Lin, T.-S. Chua, et al., Videoqa in the era of llms: An empirical study, International Journal of Computer Vision (2025) 1–24
2025
Closest in time.
T. Ning, K. Lu, X. Jiang, H. Pei, J. Xue, Dinoquery: Promoting small 3d object detection with textual prompt, IEEE Transactions on Circuits and Systems for Video Technology (2025) 1–1 doi:10.1109/TCSVT.2025.3557950
2025
Closest in time.
S. Li, T. Ji, X. Fan, L. Lu, L. Yang, Y. Yang, Z. Xi, R. Zheng, Y. Wang, xh.zhao, T. Gui, Q. Zhang, X. Huang, Have the VLMs lost confidence? a study of sycophancy in VLMs , in: The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=E2PFv7ad3p
2025
Closest in time.