Fetching the paper…
Reading the bibliography…
Knowledge-based visual question answering (VQA) requires external knowledge beyond the image to answer the question.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in NeurIPS , 2020, pp. 1877–1901
1901
Earlier work this paper cites.
H. Liu and P. Singh, “Conceptnet: a practical commonsense reasoning tool-kit,” BT technology journal , vol. 22, no. 4, pp. 211–226, 2004
2004
Earlier work this paper cites.
D. Vrandečić and M. Krötzsch, “Wikidata: A free collaborative knowledgebase,” Communications of the ACM , vol. 57, no. 10, pp. 78–85, 2014
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in NeurIPS , 2014
2014
Earlier work this paper cites.
P. Wang, Q. Wu, C. Shen, A. Dick, and A. Van Den Hengel, “Fvqa: Fact-based visual question answering,” IEEE TPAMI , vol. 40, no. 10, pp. 2413–2427, 2017
2017
Earlier work this paper cites.
P. Wang, Q. Wu, C. Shen, A. R. Dick, and A. van den Hengel, “Explicit knowledge-based reasoning for visual question answering,” in IJCAI , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” vol. 30, 2017
2017
Earlier work this paper cites.
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick, “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,” in CVPR , 2017, pp. 2901–2910
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,” in CVPR , 2017
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” IJCV , vol. 123, no. 1, pp. 32–73, 2017
2017
Earlier work this paper cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in CVPR , 2018, pp. 6077–6086
2018
Earlier work this paper cites.
J.-H. Kim, J. Jun, and B.-T. Zhang, “Bilinear attention networks,” vol. 31, 2018
2018
Earlier work this paper cites.
J. Yu, J. Li, Z. Yu, and Q. Huang, “Multimodal transformer with multi-view visual representation for image captioning,” IEEE transactions on circuits and systems for video technology , vol. 30, no. 12, pp. 4467–4480, 2019
2019
Earlier work this paper cites.
R. Hong, D. Liu, X. Mo, X. He, and H. Zhang, “Learning to compose and reason with language tree structures for visual grounding,” IEEE transactions on pattern analysis and machine intelligence , vol. 44, no. 2, pp. 684–696, 2019
2019
Earlier work this paper cites.
Z. Yu, J. Yu, Y. Cui, D. Tao, and Q. Tian, “Deep modular co-attention networks for visual question answering,” in CVPR , 2019, pp. 6281–6290
2019
Earlier work this paper cites.
H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,” EMNLP , 2019
2019
Earlier work this paper cites.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “Ok-vqa: A visual question answering benchmark requiring external knowledge,” in CVPR , 2019, pp. 3195–3204
2019
Earlier work this paper cites.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in CVPR , 2019, pp. 8317–8326
2019
Earlier work this paper cites.
L. Li, Z. Gan, Y. Cheng, and J. Liu, “Relation-aware graph attention network for visual question answering,” in ICCV , 2019, pp. 10 313–10 322
2019
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in NeurIPS , 2019
2019
Earlier work this paper cites.
D. A. Hudson and C. D. Manning, “Gqa: A new dataset for real-world visual reasoning and compositional question answering,” in CVPR , 2019, pp. 6700–6709
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, E. Valveny, C. Jawahar, and D. Karatzas, “Scene text visual question answering,” in ICCV , 2019, pp. 4291–4301
2019
Cited alongside, same era.
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei et al. , “Oscar: Object-semantics aligned pre-training for vision-language tasks,” in ECCV , 2020, pp. 121–137
2020
Cited alongside, same era.
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “Uniter: Universal image-text representation learning,” in ECCV , 2020, pp. 104–120
2020
Cited alongside, same era.
R. Hu, A. Singh, T. Darrell, and M. Rohrbach, “Iterative answer prediction with pointer-augmented multimodal transformers for textvqa,” in CVPR , 2020
2020
Cited alongside, same era.
S. Shen, L. H. Li, H. Tan, M. Bansal, A. Rohrbach, K.-W. Chang, Z. Yao, and K. Keutzer, “How much can clip benefit vision-and-language tasks?” ICLR , 2022
2022
Later among the works it cites.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” ICML , 2022
2022
Later among the works it cites.
C. Li, H. Xu, J. Tian, W. Wang, M. Yan, B. Bi, J. Ye, H. Chen, G. Xu, Z. Cao et al. , “mplug: Effective and efficient vision-language learning by cross-modal skip-connections,” pp. 7241–7259, 2022
2022
Later among the works it cites.
A. F. Biten, R. Litman, Y. Xie, S. Appalaraju, and R. Manmatha, “Latr: Layout-aware transformer for scene-text vqa,” in CVPR , 2022, pp. 16 548–16 558
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
L. Gui, B. Wang, Q. Huang, A. Hauptmann, Y. Bisk, and J. Gao, “Kat: A knowledge augmented transformer for vision-and-language,” NAACL , 2021
2021
Cited alongside, same era.
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao, “Vinvl: Revisiting visual representations in vision-language models,” in CVPR , 2021, pp. 5579–5588
2021
Cited alongside, same era.
Y. Cui, Z. Yu, C. Wang, Z. Zhao, J. Zhang, M. Wang, and J. Yu, “Rosita: Enhancing vision-and-language semantic alignments via cross-and intra-modal knowledge integration,” in ACM MM , 2021, pp. 797–806
2021
Cited alongside, same era.
Z. Yang, Y. Lu, J. Wang, X. Yin, D. Florencio, L. Wang, C. Zhang, L. Zhang, and J. Luo, “Tap: Text-aware pre-training for text-vqa and text-caption,” in CVPR , 2021, pp. 8751–8761
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021, pp. 8748–8763
2021
Cited alongside, same era.
M. Luo, Y. Zeng, P. Banerjee, and C. Baral, “Weakly-supervised visual-retriever-reader for knowledge-based question answering,” EMNLP , pp. 6417–6431, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Deng, Z. Yang, D. Liu, T. Chen, W. Zhou, Y. Zhang, H. Li, and W. Ouyang, “Transvg++: End-to-end visual grounding with language conditioned vision transformer,” IEEE transactions on pattern analysis and machine intelligence , vol. 45, no. 11, pp. 13 636–13 652, 2023
2023
Closest in time.
Z. Shao, Z. Yu, M. Wang, and J. Yu, “Prompting large language models with answer heuristics for knowledge-based visual question answering,” in CVPR , 2023, pp. 14 974–14 983
2023
Closest in time.
Y. Hu, H. Hua, Z. Yang, W. Shi, N. A. Smith, and J. Luo, “Promptcap: Prompt-guided image captioning for vqa with gpt-3,” in ICCV , 2023, pp. 2963–2975
2023
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in NeurIPS , 2023, pp. 34 892–34 916
2023
Closest in time.
2023
Closest in time.
M. AI, “Mistral 7b,” arXiv preprint arXiv:2310.06825 , 2023
2023
Closest in time.
X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer et al. , “Pali: A jointly-scaled multilingual language-image model,” in ICLR , 2023
2023
Closest in time.
W. Dai, J. Li, D. LI, A. Tiong, J. Zhao, W. Wang, B. Li, P. N. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” in NeurIPS , 2023, pp. 49 250–49 267
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei, “Voxposer: Composable 3d value maps for robotic manipulation with language models,” in Conference on Robot Learning . PMLR, 2023, pp. 540–562
2023
Closest in time.
2023
Closest in time.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in CVPR , 2024, pp. 26 296–26 306
2024
Closest in time.
2025
Closest in time.
Y. Guo, L. Nie, Y. Wong, Y. Liu, Z. Cheng, and M. Kankanhalli, “A unified end-to-end retriever-reader framework for knowledge-based vqa,” in ACM MM , 2022, pp. 2061–2069
2069
Closest in time.