Fetching the paper…
Reading the bibliography…
Contrastive decoding strategies are widely used to mitigate object hallucinations in multimodal large language models (MLLMs).
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, and et al · 2014
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
nocaps: novel object captioning at scale
Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing
Vedika Agarwal, Rakshith Shetty, and Mario Fritz · 2020
Earlier work this paper cites.
Counterfactual vqa: A cause-effect look at language bias
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen · 2021
Earlier work this paper cites.
Let there be a clock on the beach: Reducing object hallucination in image captioning
Ali Furkan Biten, Lluís Gómez, and Dimosthenis Karatzas · 2022
Earlier work this paper cites.
Swapmix: Diagnosing and regularizing the over-reliance on visual context in visual question answering
Vipul Gupta, Zhuowan Li, Adam Kortylewski, Chenyu Zhang, Yingwei Li, and Alan Yuille · 2022
Earlier work this paper cites.
Visual perturbation-aware collaborative learning for overcoming the language prior problem
Yudong Han, Liqiang Nie, Jianhua Yin, Jianlong Wu, and Yan Yan · 2022
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Xiang Lisa Li, Ari Holtzman, and et al · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Earlier work this paper cites.
A-okvqa: A benchmark for visual question answering using world knowledge
Dustin Schwenk, Apoorv Khandelwal, and et al · 2022
Earlier work this paper cites.
Yike Wu, Yu Zhao, Shiwan Zhao, Ying Zhang, Xiaojie Yuan, Guoqing Zhao, and Ning Jiang · 2022
Cited alongside, same era.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Cited alongside, same era.
Introducing our multimodal models, 2023
Rohan Bavishi, Erich Elsen, and et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang and Zhuohan et al Li · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Wenliang Dai and Junnan Li et al · 2023
Cited alongside, same era.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Overcoming language priors with self-contrastive learning for visual question answering
Hong Yan, Lijun Liu, Xupeng Feng, and Qingsong Huang · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Qinghao Ye, Haiyang Xu, and et al · 2023
Later among the works it cites.
Gpt4roi: Instruction tuning large language model on region-of-interest
Shilong Zhang, Peize Sun, and et al · 2023
Later among the works it cites.
Overcoming language priors with counterfactual inference for visual question answering
Ren Zhibo, Wang Huizhen, Zhu Muhua, Wang Yichao, Xiao Tong, and Zhu Jingbo · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chaoyou Fu, Peixian Chen, and et al · 2023
Cited alongside, same era.
Detecting and preventing hallucinations in large vision language models
Anisha Gunjal, Jihan Yin, and Erhan Bas · 2023
Cited alongside, same era.
Advancing medical imaging with language models: A journey from n-grams to chatgpt
Mingzhe Hu, Shaoyan Pan, Yuheng Li, and Xiaofeng Yang · 2023
Cited alongside, same era.
Negative object presence evaluation (nope) to measure object hallucination in vision-language models
Holy Lovenia, Wenliang Dai, Samuel Cahyawijaya, Ziwei Ji, and Pascale Fung · 2023
Cited alongside, same era.
Llm as a robotic brain: Unifying egocentric memory and control
Jinjie Mai, Jun Chen, Bing Li, Guocheng Qian, Mohamed Elhoseiny, and Bernard Ghanem · 2023
Cited alongside, same era.
Stanford alpaca: an instruction-following llama model (2023)
Rohan Taori, Ishaan Gulrajani, and et al · 2023
Cited alongside, same era.
Chatcad: Interactive computer-aided diagnosis on medical image using large language models
Sheng Wang, Zihao Zhao, Xi Ouyang, Qian Wang, and Dinggang Shen · 2023
Cited alongside, same era.
Later among the works it cites.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, and et al · 2024
Later among the works it cites.
Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation
Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu · 2024
Later among the works it cites.
Self-introspective decoding: Alleviating hallucinations for large vision-language models, 2024
Fushuo Huo, Wenchao Xu, Zhong Zhang, Haozhao Wang, Zhicheng Chen, and Peilin Zhao · 2024
Later among the works it cites.
Hallucination augmented contrastive learning for multimodal large language model
Chaoya Jiang, Haiyang Xu, and et al · 2024
Later among the works it cites.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Sicong Leng, Hang Zhang, and et al · 2024
Later among the works it cites.
Llava-next: Stronger llms supercharge multimodal capabilities in the wild, 2024
Bo Li, Kaichen Zhang, and et al · 2024
Later among the works it cites.
Introducing meta llama 3: The most capable openly available llm to date
AI Meta · 2024
Later among the works it cites.
Mitigating hallucinations in large vision-language models with instruction contrastive decoding
Xintong Wang, Jingheng Pan, and et al · 2024
Later among the works it cites.