Fetching the paper…
Reading the bibliography…
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in processing both visual and textual information.
Language models are few-shot learners
Tom B Brown. 2020 · 2005
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, Andrew Y Ng, et al. 2011 · 2011
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky. 2015 · 2015
Earlier work this paper cites.
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018 · 2018
Earlier work this paper cites.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aaron Van Den Oord, Alex Alemi, and George Tucker. 2019 · 2019
Earlier work this paper cites.
Addressing the polysemy problem in language modeling with attentional multi-sense embeddings
Rao Ma, Lesheng Jin, Qi Liu, Lu Chen, and Kai Yu. 2020 · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Earlier work this paper cites.
Plausible may not be faithful: Probing object hallucination in vision-language pre-training
Wenliang Dai, Zihan Liu, Ziwei Ji, Dan Su, and Pascale Fung. 2022 · 2022
Earlier work this paper cites.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Victor Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Y Zou. 2022 · 2022
Earlier work this paper cites.
Vision-language pre-training with triple contrastive learning
Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, and Junzhou Huang. 2022 · 2022
Earlier work this paper cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, et al. 2023 · 2023
Earlier work this paper cites.
Ciem: Contrastive instruction evaluation method for better instruction tuning
Hongyu Hu, Jiyuan Zhang, Minyi Zhao, and Zhenbang Sun. 2023 · 2023
Earlier work this paper cites.
Volcano: mitigating multimodal hallucination through self-feedback guided revision
Seongyun Lee, Sue Hyun Park, Yongrae Jo, and Minjoon Seo. 2023 · 2023
Earlier work this paper cites.
Mitigating hallucination in large multi-modal models via robust instruction tuning
Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. 2023 · 2023
Earlier work this paper cites.
Exposing and addressing cross-task inconsistency in unified vision-language models
Adyasha Maharana, Amita Kamath, Christopher Clark, Mohit Bansal, and Aniruddha Kembhavi. 2023 · 2023
Earlier work this paper cites.
Chatgpt can now see, hear, and speak
OpenAI · 2023
Earlier work this paper cites.
Aligning large multimodal models with factually augmented rlhf
Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al. 2023 · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023 · 2023
Earlier work this paper cites.
Evaluation and analysis of hallucination in large vision-language models
Junyang Wang, Yiyang Zhou, Guohai Xu, Pengcheng Shi, Chenlin Zhao, Haiyang Xu, Qinghao Ye, Ming Yan, Ji Zhang, Jihua Zhu, et al. 2023 · 2023
Earlier work this paper cites.
Woodpecker: Hallucination correction for multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Tong Xu, Hao Wang, Dianbo Sui, Yunhang Shen, Ke Li, Xing Sun, and Enhong Chen. 2023 · 2023
Earlier work this paper cites.
Ferret: Refer and ground anything anywhere at any granularity
Haoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du, Bowen Zhang, Zirui Wang, Liangliang Cao, Shih-Fu Chang, and Yinfei Yang. 2023 · 2023
Earlier work this paper cites.
Bohan Zhai, Shijia Yang, Xiangchen Zhao, Chenfeng Xu, Sheng Shen, Dongdi Zhao, Kurt Keutzer, Manling Li, Tan Yan, and Xiangjun Fan. 2023 · 2023
Earlier work this paper cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Earlier work this paper cites.
Wenbin An, Feng Tian, Sicong Leng, Jiahao Nie, Haonan Lin, QianYing Wang, Guang Dai, Ping Chen, and Shijian Lu. 2024 · 2024
Cited alongside, same era.
Claude 3.5 sonnet
Anthropic. 2024 · 2024
Cited alongside, same era.
Hallucination of multimodal large language models: A survey
Zechen Bai, Pichao Wang, Tianjun Xiao, Tong He, Zongbo Han, Zheng Zhang, and Mike Zheng Shou. 2024 · 2024
Cited alongside, same era.
Videocon: Robust video-language alignment via contrast captions
Hritik Bansal, Yonatan Bitton, Idan Szpektor, Kai-Wei Chang, and Aditya Grover. 2024 · 2024
Cited alongside, same era.
Mitigating open-vocabulary caption hallucinations
Assaf Ben-Kish, Moran Yanuka, Morris Alper, Raja Giryes, and Hadar Averbuch-Elor. 2024 · 2024
Cited alongside, same era.
Mitigating hallucinations in large vision-language models via summary-guided decoding
Kyungmin Min, Minbeom Kim, Kang-il Lee, Dongryeol Lee, and Kyomin Jung. 2024 · 2024
Later among the works it cites.
Towards interpreting visual information processing in vision-language models
Clement Neo, Luke Ong, Philip Torr, Mor Geva, David Krueger, and Fazl Barez. 2024 · 2024
Later among the works it cites.
Mitigating dialogue hallucination for large multi-modal models via adversarial instruction tuning
Dongmin Park, Zhaofang Qian, Guangxing Han, and Ser-Nam Lim. 2024 · 2024
Later among the works it cites.
Alleviating hallucination in large vision-language models with active retrieval augmentation
Xiaoye Qu, Qiyuan Chen, Wei Wei, Jishuo Sun, and Jianfeng Dong. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mafa: Managing false negatives for vision-language pre-training
Jaeseok Byun, Dohoon Kim, and Taesup Moon. 2024 · 2024
Cited alongside, same era.
Visually dehallucinative instruction generation
Sungguk Cha, Jusung Lee, Younghyun Lee, and Cheoljong Yang. 2024 · 2024
Cited alongside, same era.
Ailin Deng, Zhirui Chen, and Bryan Hooi. 2024 · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Cited alongside, same era.
Multi-modal hallucination control by visual information grounding
Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, and Stefano Soatto. 2024 · 2024
Cited alongside, same era.
Do more details always introduce more hallucinations in lvlm-based image captioning?
Mingqian Feng, Yunlong Tang, Zeliang Zhang, and Chenliang Xu. 2024 · 2024
Cited alongside, same era.
Metatoken: Detecting hallucination in image descriptions by meta classification
Laura Fieback, Jakob Spiegelberg, and Hanno Gottschalk. 2024 · 2024
Cited alongside, same era.
Pranab Sahoo, Prabhash Meharia, Akash Ghosh, Sriparna Saha, Vinija Jain, and Aman Chadha. 2024 · 2024
Later among the works it cites.
Towards retrieval-augmented architectures for image captioning
Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Alessandro Nicolosi, and Rita Cucchiara. 2024 · 2024
Later among the works it cites.
From pixels to tokens: Revisiting object hallucinations in large vision-language models
Yuying Shang, Xinyi Zeng, Yutao Zhu, Xiao Yang, Zhengwei Fang, Jingyuan Zhang, Jiawei Chen, Zinan Liu, and Yu Tian. 2024 · 2024
Later among the works it cites.
A vision check-up for language models
Pratyusha Sharma, Tamar Rott Shaham, Manel Baradad, Stephanie Fu, Adrian Rodriguez-Munoz, Shivam Duggal, Phillip Isola, and Antonio Torralba. 2024 · 2024
Later among the works it cites.
Assessment of multimodal large language models in alignment with human values
Zhelun Shi, Zhipin Wang, Hongxing Fan, Zaibin Zhang, Lijun Li, Yongting Zhang, Zhenfei Yin, Lu Sheng, Yu Qiao, and Jing Shao. 2024 · 2024
Later among the works it cites.
Exploring the alignment landscape: Llms and geometric deep models in protein representation
Dong Shu, Bingbing Duan, Kai Guo, Kaixiong Zhou, Jiliang Tang, and Mengnan Du. 2024 · 2024
Later among the works it cites.
Textsquare: Scaling up text-centric visual instruction tuning
Jingqun Tang, Chunhui Lin, Zhen Zhao, Shu Wei, Binghong Wu, Qi Liu, Hao Feng, Yang Li, Siqi Wang, Lei Liao, et al. 2024 · 2024
Later among the works it cites.
Detecting and mitigating hallucination in large vision language models via fine-grained ai feedback
Wenyi Xiao, Ziwei Huang, Leilei Gan, Wanggui He, Haoyuan Li, Zhelun Yu, Hao Jiang, Fei Wu, and Linchao Zhu. 2024 · 2024
Later among the works it cites.
Yuxi Xie, Guanzhen Li, Xiao Xu, and Min-Yen Kan. 2024 · 2024
Later among the works it cites.
Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models
Peng Xu, Wenqi Shao, Kaipeng Zhang, Peng Gao, Shuo Liu, Meng Lei, Fanqing Meng, Siyuan Huang, Yu Qiao, and Ping Luo. 2024 · 2024
Later among the works it cites.
List items one by one: A new data source and learning paradigm for multimodal llms
An Yan, Zhengyuan Yang, Junda Wu, Wanrong Zhu, Jianwei Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Julian McAuley, Jianfeng Gao, et al. 2024 · 2024
Later among the works it cites.
Less is more: Mitigating multimodal hallucination from an eos decision perspective
Zihao Yue, Liang Zhang, and Qin Jin. 2024 · 2024
Later among the works it cites.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Simon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan, Peter Tong, Yifei Zhou, Alane Suhr, Saining Xie, Yann LeCun, Yi Ma, et al. 2024 · 2024
Later among the works it cites.
Rethinking misalignment in vision-language model adaptation from a causal perspective
Yanan Zhang, Jiangmeng Li, Lixiang Liu, and Wenwen Qiang. 2024 · 2024
Later among the works it cites.
Weihong Zhong, Xiaocheng Feng, Liang Zhao, Qiming Li, Lei Huang, Yuxuan Gu, Weitao Ma, Yuan Xu, and Bing Qin. 2024 · 2024
Later among the works it cites.
Aligning modalities in vision large language models via preference fine-tuning
Yiyang Zhou, Chenhang Cui, Rafael Rafailov, Chelsea Finn, and Huaxiu Yao. 2024 · 2024
Later among the works it cites.
An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. 2025 · 2025
Closest in time.
Clip-dpo: Vision-language models as a source of preference for fixing hallucinations in lvlms
Yassine Ouali, Adrian Bulat, Brais Martinez, and Georgios Tzimiropoulos. 2025 · 2025
Closest in time.
Reflective instruction tuning: Mitigating hallucinations in large vision-language models
Jinrui Zhang, Teng Wang, Haigang Zhang, Ping Lu, and Feng Zheng. 2025 · 2025
Closest in time.
Can hallucination correction improve video-language alignment?
Lingjun Zhao, Mingyang Xie, Paola Cascante-Bonilla, Hal Daumé III, and Kwonjoon Lee. 2025 · 2025
Closest in time.