Fetching the paper…
Reading the bibliography…
Visual question answering (VQA) is a task where an image is given, and a series of questions are asked about the image.
Vizwiz: nearly real-time answers to visual questions
J. P. Bigham, C. Jayant, H. Ji, G. Little, A. Miller, R. C. Miller, R. Miller, A. Tatarowicz, B. White, S. White, et al · 2010
Earlier work this paper cites.
Compositional attention networks for machine reasoning
D. A. Hudson and C. D. Manning · 2018
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
J. J. Lau, S. Gayen, A. Ben Abacha, and D. Demner-Fushman · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Earlier work this paper cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion · 2021
Earlier work this paper cites.
Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering
B. Liu, L.-M. Zhan, L. Xu, L. Ma, Y. Yang, and X.-M. Wu · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Multimodal few-shot learning with frozen language models
M. Tsimpoukelli, J. L. Menick, S. Cabi, S. Eslami, O. Vinyals, and F. Hill · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Cited alongside, same era.
Plug-and-play vqa: Zero-shot vqa by conjoining large pretrained models with zero training
A. M. H. Tiong, J. Li, B. Li, S. Savarese, and S. C. Hoi · 2022
Cited alongside, same era.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang · 2022
Cited alongside, same era.
An empirical study of gpt-3 for few-shot knowledge-based vqa
Z. Yang, Z. Gan, J. Wang, X. Hu, Y. Lu, Z. Liu, and L. Wang · 2022
Cited alongside, same era.
Coca: Contrastive captioners are image-text foundation models
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu · 2022
A comprehensive study of gpt-4v’s multimodal capabilities in medical imaging
Y. Li, Y. Liu, Z. Wang, X. Liang, L. Liu, L. Wang, L. Cui, Z. Tu, L. Wang, and L. Zhou · 2023
Later among the works it cites.
Medical visual question answering: A survey
Z. Lin, D. Zhang, Q. Tao, D. Shi, G. Haffari, Q. Wu, M. He, and Z. Ge · 2023
Later among the works it cites.
Med-flamingo: a multimodal medical few-shot learner
M. Moor, Q. Huang, S. Wu, M. Yasunaga, Y. Dalmia, J. Leskovec, C. Zakka, E. P. Reis, and P. Rajpurkar · 2023
Later among the works it cites.
Idealgpt: Iteratively decomposing vision and language reasoning via large language models
H. You, R. Sun, Z. Wang, L. Chen, G. Wang, H. Ayyubi, K.-W. Chang, and S.-F. Chang · 2023
Later among the works it cites.
Vicor: Bridging visual understanding and commonsense reasoning with large language models
K. Zhou, K. Lee, T. Misu, and X. E. Wang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
From images to textual prompts: Zero-shot visual question answering with frozen large language models
J. Guo, J. Li, D. Li, A. M. H. Tiong, B. Li, D. Tao, and S. Hoi · 2023
Cited alongside, same era.
Expert knowledge-aware image difference graph representation learning for difference-aware medical visual question answering
X. Hu, L. Gu, Q. An, M. Zhang, L. Liu, K. Kobayashi, T. Harada, R. M. Summers, and Y. Zhu · 2023
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick
Cited in the paper.
Inferring and executing programs for visual reasoning
J. Johnson, B. Hariharan, L. Van Der Maaten, J. Hoffman, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick
Cited in the paper.
K. Zhang, J. Yu, Z. Yan, Y. Liu, E. Adhikarla, S. Fu, X. Chen, C. Chen, Y. Zhou, X. Li, et al
Cited in the paper.
Pmc-vqa: Visual instruction tuning for medical visual question answering
X. Zhang, C. Wu, Z. Zhao, W. Lin, Y. Zhang, Y. Wang, and W. Xie
Cited in the paper.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Closest in time.
Improving automatic vqa evaluation using large language models
O. Mañas, B. Krojer, and A. Agrawal · 2024
Closest in time.