Fetching the paper…
Reading the bibliography…
Large Vision-Language Models (LVLMs) excel in integrating visual and linguistic contexts to produce detailed content, facilitating applications such as image captioning.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017 · 2017
Earlier work this paper cites.
A survey on automatic image caption generation
Shuang Bai and Shan An. 2018 · 2018
Earlier work this paper cites.
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018 · 2018
Earlier work this paper cites.
An overview of image caption generation methods
Haoran Wang, Yue Zhang, and Xiaosheng Yu. 2020 · 2020
Earlier work this paper cites.
A course-focused dual curriculum for image captioning
Mohammad Alsharid, Rasheed El-Bouri, Harshita Sharma, Lior Drukker, Aris T. Papageorghiou, and J. Alison Noble. 2021 · 2021
Earlier work this paper cites.
Gaze-assisted automatic captioning of fetal ultrasound videos using three-way multi-modal deep neural networks
Mohammad Alsharid, Yifan Cai, Harshita Sharma, Lior Drukker, Aris T. Papageorghiou, and J. Alison Noble. 2022 · 2022
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2022 · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022 · 2022
Cited alongside, same era.
Vision-and-language pretrained models: A survey
Siqu Long, Feiqi Cao, Soyeon Caren Han, and Haiqin Yang. 2022 · 2022
Cited alongside, same era.
Advancing medical imaging with language models: A journey from n-grams to chatgpt
Mingzhe Hu, Shaoyan Pan, Yuheng Li, and Xiaofeng Yang. 2023 · 2023
Cited alongside, same era.
Adapt: Action-aware driving caption transformer
Bu Jin, Xinyu Liu, Yupeng Zheng, Pengfei Li, Hao Zhao, Tong Zhang, Yuhang Zheng, Guyue Zhou, and Jingjing Liu. 2023 · 2023
Cited alongside, same era.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Analyzing and mitigating object hallucination in large vision-language models
Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, and Huaxiu Yao. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Later among the works it cites.
Ailin Deng, Zhirui Chen, and Bryan Hooi. 2024 · 2024
Closest in time.
Multi-modal hallucination control by visual information grounding
Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, and Stefano Soatto. 2024 · 2024
Closest in time.
Detecting and preventing hallucinations in large vision language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing. 2023 · 2023
Cited alongside, same era.
Evaluation and mitigation of agnosia in multimodal large language models
Jiaying Lu, Jinmeng Rao, Kezhen Chen, Xiaoyuan Guo, Yawen Zhang, Baochen Sun, Carl Yang, and Jie Yang. 2023 · 2023
Cited alongside, same era.
Caption anything: Interactive image description with diverse multimodal controls
Teng Wang, Jinrui Zhang, Junjie Fei, Yixiao Ge, Hao Zheng, Yunlong Tang, Zhe Li, Mingqi Gao, Shanshan Zhao, Ying Shan, et al. 2023 · 2023
Cited alongside, same era.
Pink: Unveiling the power of referential comprehension for multi-modal llms
Shiyu Xuan, Qingpei Guo, Ming Yang, and Shiliang Zhang. 2023 · 2023
Cited alongside, same era.
Minigpt-v2: large language model as a unified interface for vision-language multi-task learning
Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny. 2023a
Cited in the paper.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. 2023b
Cited in the paper.
Driving with llms: Fusing object-level vector modality for explainable autonomous driving
Long Chen, Oleg Sinavski, Jan Hünermann, Alice Karnsund, Andrew James Willmott, Danny Birch, Daniel Maund, and Jamie Shotton. 2023c
Cited in the paper.
Aligning large multi-modal model with robust instruction tuning
Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. 2023a
Cited in the paper.
Anisha Gunjal, Jihan Yin, and Erhan Bas. 2024 · 2024
Closest in time.
Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu. 2024 · 2024
Closest in time.
OpenAI. 2024 · 2024
Closest in time.
Mitigating dialogue hallucination for large multi-modal models via adversarial instruction tuning
Dongmin Park, Zhaofang Qian, Guangxing Han, and Ser-Nam Lim. 2024 · 2024
Closest in time.