Fetching the paper…
Reading the bibliography…
Recently, to comprehensively improve Vision Language Models (VLMs) for Visual Question Answering (VQA), several methods have been proposed to further reinforce the inference capabilities of VLMs to independently tackle VQA tasks rather than some methods that only utilize VLMs as aids to Large Language Models (LLMs).
Wikidata: a free collaborative knowledgebase
Denny Vrandečić and Markus Krötzsch · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Learning compact hash codes for multimodal representations using orthogonal deep structure
Daixin Wang, Peng Cui, Mingdong Ou, and Wenwu Zhu · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Out of the box: Reasoning with graph convolution nets for factual visual question answering
Medhini Narasimhan, Svetlana Lazebnik, and Alexander Schwing · 2018
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman · 2018
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Aishwarya Agrawal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2019
Earlier work this paper cites.
Squinting at vqa models: Introspecting vqa models with sub-questions
Ramprasaath R Selvaraju, Purva Tendulkar, Devi Parikh, Eric Horvitz, Marco Tulio Ribeiro, Besmira Nushi, and Ece Kamar · 2020
Earlier work this paper cites.
Conceptbert: Concept-aware representation for visual question answering
François Gardères, Maryam Ziaeefard, Baptiste Abeloos, and Freddy Lecue · 2020
Earlier work this paper cites.
Aser: A large-scale eventuality knowledge graph
Hongming Zhang, Xin Liu, Haojie Pan, Yangqiu Song, and Cane Wing-Ki Leung · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Understanding guided image captioning performance across domains
Edwin G. Ng, Bo Pang, Piyush Sharma, and Radu Soricut · 2020
Earlier work this paper cites.
Just ask: Learning to answer questions from millions of narrated videos
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2021
Earlier work this paper cites.
Ruart: A novel text-centered solution for text-based visual question answering
Zan-Xia Jin, Heran Wu, Chun Yang, Fang Zhou, Jingyan Qin, Lei Xiao, and Xu-Cheng Yin · 2021
Earlier work this paper cites.
X-ggm: Graph generative modeling for out-of-distribution generalization in visual question answering
Jingjing Jiang, Ziyi Liu, Yifan Liu, Zhixiong Nan, and Nanning Zheng · 2021
Earlier work this paper cites.
Focal and composed vision-semantic modeling for visual question answering
Yudong Han, Yangyang Guo, Jianhua Yin, Meng Liu, Yupeng Hu, and Liqiang Nie · 2021
Earlier work this paper cites.
Krisp: Integrating implicit and symbolic knowledge for open-domain knowledge-based vqa
Kenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta, and Marcus Rohrbach · 2021
Earlier work this paper cites.
Modeling voters in multi-winner approval voting
Jaelle Scheuerman, Jason Harman, Nicholas Mattei, and K Brent Venable · 2021
Earlier work this paper cites.
Coarse-to-fine reasoning for visual question answering
Binh X Nguyen, Tuong Do, Huy Tran, Erman Tjiputra, Quang D Tran, and Anh Nguyen · 2022
Earlier work this paper cites.
Positional attention guided transformer-like architecture for visual question answering
Aihua Mao, Zhi Yang, Ken Lin, Jun Xuan, and Yong-Jin Liu · 2022
Earlier work this paper cites.
A unified end-to-end retriever-reader framework for knowledge-based vqa
Yangyang Guo, Liqiang Nie, Yongkang Wong, Yibing Liu, Zhiyong Cheng, and Mohan Kankanhalli · 2022
Earlier work this paper cites.
Ai-vqa: Visual question answering based on agent interaction with interpretability
Rengang Li, Cong Xu, Zhenhua Guo, Baoyu Fan, Runze Zhang, Wei Liu, Yaqian Zhao, Weifeng Gong, and Endong Wang · 2022
Earlier work this paper cites.
An empirical study of gpt-3 for few-shot knowledge-based vqa
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang · 2022
Earlier work this paper cites.
Promptcap: Prompt-guided task-aware image captioning
Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi, Noah A Smith, and Jiebo Luo · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Cited alongside, same era.
Unsupervised and pseudo-supervised vision-language alignment in visual dialog
Feilong Chen, Duzhen Zhang, Xiuyi Chen, Jing Shi, Shuang Xu, and Bo Xu · 2022
Cited alongside, same era.
Learning efficient multi-agent cooperative visual exploration
Chao Yu, Xinyi Yang, Jiaxuan Gao, Huazhong Yang, Yu Wang, and Yi Wu · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Cited alongside, same era.
Woodpecker: Hallucination correction for multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Tong Xu, Hao Wang, Dianbo Sui, Yunhang Shen, Ke Li, Xing Sun, and Enhong Chen · 2023
Closest in time.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Closest in time.
Cogvlm: Visual expert for pretrained language models
Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, et al · 2023
Closest in time.
Clip-vg: Self-paced curriculum adapting of clip for visual grounding
Linhui Xiao, Xiaoshan Yang, Fang Peng, Ming Yan, Yaowei Wang, and Changsheng Xu · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A-okvqa: A benchmark for visual question answering using world knowledge
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi · 2022
Cited alongside, same era.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Self-pt: Adaptive self-prompt tuning for low-resource visual question answering
Bowen Yuan, Sisi You, and Bing-Kun Bao · 2023
Cited alongside, same era.
Measuring and improving chain-of-thought reasoning in vision-language models
Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, and Ajay Divakaran · 2023
Cited alongside, same era.
Idealgpt: Iteratively decomposing vision and language reasoning via large language models
Haoxuan You, Rui Sun, Zhecan Wang, Long Chen, Gengyu Wang, Hammad A Ayyubi, Kai-Wei Chang, and Shih-Fu Chang · 2023
Cited alongside, same era.
Improving zero-shot visual question answering via large language models with reasoning question prompts
Yunshi Lan, Xiang Li, Xin Liu, Yang Li, Wei Qin, and Weining Qian · 2023
Cited alongside, same era.
Closest in time.
Llava-med: Training a large language-and-vision assistant for biomedicine in one day
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao · 2023
Closest in time.
Not all features matter: Enhancing few-shot clip with adaptive prior refinement
Xiangyang Zhu, Renrui Zhang, Bowei He, Aojun Zhou, Dong Wang, Bin Zhao, and Peng Gao · 2023
Closest in time.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi · 2023
Closest in time.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Closest in time.
Retrieval-based knowledge augmented vision language pre-training
Jiahua Rao, Zifei Shan, Longpo Liu, Yao Zhou, and Yuedong Yang · 2023
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al · 2023
Closest in time.
Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, and Heng Ji · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Closest in time.
A survey of efficient fine-tuning methods for vision-language models — prompt and adapter
Jialu Xing, Jianping Liu, Jian Wang, Lulu Sun, Xi Chen, Xunxun Gu, and Yingfei Wang · 2024
Closest in time.
Exploring question decomposition for zero-shot vqa
Zaid Khan, Vijay Kumar BG, Samuel Schulter, Manmohan Chandraker, and Yun Fu · 2024
Closest in time.
Haibi Wang and Weifeng Ge · 2024
Closest in time.
Prompting large language models with fine-grained visual relations from scene graph for visual question answering
Jiapeng Liu, Chengyang Fang, Liang Li, Bing Li, Dayong Hu, and Can Ma · 2024
Closest in time.
What is missing in multilingual visual reasoning and how to fix it
Yueqi Song, Simran Khanuja, and Graham Neubig · 2024
Closest in time.
Visual question decomposition on multimodal large language models
Haowei Zhang, Jianzhe Liu, Zhen Han, Shuo Chen, Bailan He, Volker Tresp, Zhiqiang Xu, and Jindong Gu · 2024
Closest in time.
Cross modality bias in visual question answering: A causal view with possible worlds vqa
Ali Vosoughi, Shijian Deng, Songyang Zhang, Yapeng Tian, Chenliang Xu, and Jiebo Luo · 2024
Closest in time.
Clova: A closed-loop visual assistant with tool usage and update
Zhi Gao, Yuntao Du, Xintong Zhang, Xiaojian Ma, Wenjuan Han, Song-Chun Zhu, and Qing Li · 2024
Closest in time.
More agents is all you need
junyou li, Qin Zhang, Yangbin Yu, QIANG FU, and Deheng Ye · 2024
Closest in time.
Bliva: A simple multimodal llm for better handling of text-rich visual questions
Wenbo Hu, Yifan Xu, Yi Li, Weiyue Li, Zeyuan Chen, and Zhuowen Tu · 2024
Closest in time.
Cogagent: A visual language model for gui agents
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, et al · 2024
Closest in time.