Fetching the paper…
Reading the bibliography…
Large Vision-Language Models (LVLMs) have made remarkable strides in multimodal tasks such as visual question answering, visual grounding, and complex reasoning.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020 · 2020
Earlier work this paper cites.
Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval
Tobias Weyand, Andre Araujo, Bingyi Cao, and Jack Sim. 2020 · 2020
Earlier work this paper cites.
Benchmarking representation learning for natural world image collections
Grant Van Horn, Elijah Cole, Sara Beery, Kimberly Wilber, Serge Belongie, and Oisin Mac Aodha. 2021 · 2021
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023 · 2023
Earlier work this paper cites.
Can pre-trained vision and language models answer visual information-seeking questions?
Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, and Ming-Wei Chang. 2023 · 2023
Earlier work this paper cites.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023 · 2023
Earlier work this paper cites.
Interactive task planning with language models
Boyi Li, Philipp Wu, Pieter Abbeel, and Jitendra Malik. 2023 · 2023
Earlier work this paper cites.
Large language model is not a good few-shot information extractor, but a good reranker for hard samples!
Yubo Ma, Yixin Cao, Yong Hong, and Aixin Sun. 2023 · 2023
Earlier work this paper cites.
Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories
Thomas Mensink, Jasper Uijlings, Lluis Castrejon, Arushi Goel, Felipe Cadar, Howard Zhou, Fei Sha, André Araujo, and Vittorio Ferrari. 2023 · 2023
Earlier work this paper cites.
Large language models are effective text rankers with pairwise ranking prompting
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, et al. 2023 · 2023
Earlier work this paper cites.
Eva-clip: Improved training techniques for clip at scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. 2023 · 2023
Earlier work this paper cites.
Improving visual grounding by encouraging consistent gradient-based explanations
Ziyan Yang, Kushal Kafle, Franck Dernoncourt, and Vicente Ordonez. 2023 · 2023
Earlier work this paper cites.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2023 · 2023
Earlier work this paper cites.
Hallucination of multimodal large language models: A survey
Zechen Bai, Pichao Wang, Tianjun Xiao, Tong He, Zongbo Han, Zheng Zhang, and Mike Zheng Shou. 2024 · 2024
Earlier work this paper cites.
Wiki-llava: Hierarchical retrieval-augmented generation for multimodal llms
Davide Caffagni, Federico Cocchi, Nicholas Moratelli, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2024 · 2024
Earlier work this paper cites.
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024 · 2024
Earlier work this paper cites.
Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Jiaqi Wang, et al. 2024 · 2024
Earlier work this paper cites.
Multi-modal hallucination control by visual information grounding
Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, and Stefano Soatto. 2024 · 2024
Earlier work this paper cites.
FIRST: Faster improved listwise reranking with single token decoding
Revanth Gangi Reddy, JaeHyeok Doo, Yifei Xu, Md Arafat Sultan, Deevya Swain, Avirup Sil, and Heng Ji. 2024 · 2024
Earlier work this paper cites.
Physically grounded vision-language models for robotic manipulation
Jensen Gao, Bidipta Sarkar, Fei Xia, Ted Xiao, Jiajun Wu, Brian Ichter, Anirudha Majumdar, and Dorsa Sadigh. 2024 · 2024
Cited alongside, same era.
E5-v: Universal embeddings with multimodal large language models
Ting Jiang, Minghui Song, Zihan Zhang, Haizhen Huang, Weiwei Deng, Feng Sun, Qi Zhang, Deqing Wang, and Fuzhen Zhuang. 2024 · 2024
Cited alongside, same era.
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al. 2024 · 2024
Cited alongside, same era.
Can feedback enhance semantic grounding in large vision-language models?
Yuan-Hong Liao, Rafid Mahmood, Sanja Fidler, and David Acuna. 2024 · 2024
Cited alongside, same era.
Mm-embed: Universal multimodal retrieval with multimodal llms
EchoSight: Advancing visual-language models with Wiki knowledge
Yibin Yan and Weidi Xie. 2024 · 2024
Later among the works it cites.
Guiding long-horizon task and motion planning with vision language models
Zhutian Yang, Caelan Garrett, Dieter Fox, Tomás Lozano-Pérez, and Leslie Pack Kaelbling. 2024 · 2024
Later among the works it cites.
Rankrag: Unifying context ranking with retrieval-augmented generation in llms
Yue Yu, Wei Ping, Zihan Liu, Boxin Wang, Jiaxuan You, Chao Zhang, Mohammad Shoeybi, and Bryan Catanzaro. 2024 · 2024
Later among the works it cites.
Jianhao Yuan, Shuyang Sun, Daniel Omeiza, Bo Zhao, Paul Newman, Lars Kunze, and Matthew Gadd. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin, Bryan Catanzaro, and Wei Ping. 2024 · 2024
Cited alongside, same era.
Lost in the middle: How language models use long contexts
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024 · 2024
Cited alongside, same era.
Ovis: Structural embedding alignment for multimodal large language model
Shiyin Lu, Yang Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, and Han-Jia Ye. 2024 · 2024
Cited alongside, same era.
Trust but verify: Programmatic vlm evaluation in the wild
Viraj Prabhu, Senthil Purushwalkam, An Yan, Caiming Xiong, and Ran Xu. 2024 · 2024
Cited alongside, same era.
Beyond text: Optimizing rag with multimodal inputs for industrial applications
Monica Riedler and Stefan Langer. 2024 · 2024
Cited alongside, same era.
A comprehensive survey of hallucination in large language, image, video and audio foundation models
Pranab Sahoo, Prabhash Meharia, Akash Ghosh, Sriparna Saha, Vinija Jain, and Aman Chadha. 2024 · 2024
Cited alongside, same era.
Pelican: Correcting hallucination in vision-LLMs via claim decomposition and program of thought verification
Pritish Sahu, Karan Sikka, and Ajay Divakaran. 2024 · 2024
Cited alongside, same era.
Neelabh Sinha, Vinija Jain, and Aman Chadha. 2024 · 2024
Cited alongside, same era.
Zhaxizhuoma Zhaxizhuoma, Pengan Chen, Ziniu Wu, Jiawei Sun, Dong Wang, Peng Zhou, Nieqing Cao, Yan Ding, Bin Zhao, and Xuelong Li. 2024 · 2024
Later among the works it cites.
Megapairs: Massive data synthesis for universal multimodal retrieval
Junjie Zhou, Zheng Liu, Ze Liu, Shitao Xiao, Yueze Wang, Bo Zhao, Chen Jason Zhang, Defu Lian, and Yongping Xiong. 2024 · 2024
Later among the works it cites.
Unraveling cross-modality knowledge conflicts in large vision-language models
Tinghui Zhu, Qin Liu, Fei Wang, Zhengzhong Tu, and Muhao Chen. 2024 · 2024
Later among the works it cites.
A setwise approach for effective and highly efficient zero-shot ranking with large language models
Shengyao Zhuang, Honglei Zhuang, Bevan Koopman, and Guido Zuccon. 2024 · 2024
Later among the works it cites.
Ask in any modality: A comprehensive survey on multimodal retrieval-augmented generation
Mohammad Mahdi Abootorabi, Amirhosein Zobeiri, Mahdi Dehghani, Mohammadali Mohammadkhani, Bardia Mohammadi, Omid Ghahroodi, Mahdieh Soleymani Baghshah, and Ehsaneddin Asgari. 2025 · 2025
Closest in time.
Vision-language models struggle to align entities across modalities
Iñigo Alonso, Ander Salaberria, Gorka Azkune, Jeremy Barnes, and Oier Lopez de Lacalle. 2025 · 2025
Closest in time.
Multimodal fact-checking with vision language models: A probing classifier based solution with embedding strategies
Recep Firat Cekinel, Pinar Karagoz, and Çağrı Çöltekin. 2025 · 2025
Closest in time.
Physbench: Benchmarking and enhancing vision-language models for physical world understanding
Wei Chow, Jiageng Mao, Boyi Li, Daniel Seita, Vitor Guizilini, and Yue Wang. 2025 · 2025
Closest in time.
Llm4ranking: An easy-to-use framework of utilizing large language models for document reranking
Qi Liu, Haozhe Duan, Yiqun Chen, Quanfeng Lu, Weiwei Sun, and Jiaxin Mao. 2025 · 2025
Closest in time.
A survey of multimodal retrieval-augmented generation
Lang Mei, Siyu Mo, Zhihan Yang, and Chong Chen. 2025 · 2025
Closest in time.
Re-ranking the context for multimodal retrieval augmented generation
Matin Mortaheb, Mohammad A Amir Khojastepour, Srimat T Chakradhar, and Sennur Ulukus. 2025 · 2025
Closest in time.
Defining and quantifying visual hallucinations in vision-language models
Vipula Rawte, Aryan Mishra, Amit Sheth, and Amitava Das. 2025 · 2025
Closest in time.
Self-calibrated listwise reranking with large language models
Ruiyang Ren, Yuhao Wang, Kun Zhou, Wayne Xin Zhao, Wenjie Wang, Jing Liu, Ji-Rong Wen, and Tat-Seng Chua. 2025 · 2025
Closest in time.
Learning visual grounding from generative vision and language model
Shijie Wang, Dahun Kim, Ali Taalimi, Chen Sun, and Weicheng Kuo. 2025 · 2025
Closest in time.
Physvlm: Enabling visual language models to understand robotic physical reachability
Weijie Zhou, Manli Tao, Chaoyang Zhao, Haiyun Guo, Honghui Dong, Ming Tang, and Jinqiao Wang. 2025 · 2025
Closest in time.
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Yuchen Duan, Hao Tian, Weijie Su, Jie Shao, et al. 2025 · 2025
Closest in time.