Fetching the paper…
Reading the bibliography…
With the breakthrough of multi-modal large language models, answering complex visual questions that demand advanced reasoning abilities and world knowledge has become a much more important testbed for developing AI models than ever.
Van der Maaten, L., Hinton, G.: Visualizing data using t-sne. Journal of machine learning research 9
2008
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: common objects in context. In: Eur. Conf. Comput. Vis. (2014)
2014
Earlier work this paper cites.
Vrandečić, D., Krötzsch, M.: Wikidata: a free collaborative knowledgebase. Communications of the ACM 57
2014
Earlier work this paper cites.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: Vqa: visual question answering. In: Int. Conf. Comput. Vis. (2015)
2015
Earlier work this paper cites.
Geman, D., Geman, S., Hallonquist, N., Younes, L.: Visual turing test for computer vision systems. Proceedings of the National Academy of Sciences 112
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D.: Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. (2017)
2017
Earlier work this paper cites.
Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., Girshick, R.: Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 2901–2910 (2017)
2017
Earlier work this paper cites.
Speer, R., Chin, J., Havasi, C.: Conceptnet 5.5: An open multilingual graph of general knowledge. In: AAAI (2017)
2017
Earlier work this paper cites.
Wang, P., Wu, Q., Shen, C., Dick, A., Van Den Hengel, A.: Fvqa: Fact-based visual question answering. IEEE transactions on pattern analysis and machine intelligence 40
2017
Earlier work this paper cites.
Agrawal, A., Batra, D., Parikh, D., Kembhavi, A.: Don’t just assume; look and answer: Overcoming priors for visual question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding. In: NAACL (2019)
2019
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: Int. Conf. Learn. Represent. (2019)
2019
Earlier work this paper cites.
Lu, J., Batra, D., Parikh, D., Lee, S.: Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In: Adv. Neural Inform. Process. Syst. (2019)
2019
Earlier work this paper cites.
Marino, K., Rastegari, M., Farhadi, A., Mottaghi, R.: Ok-vqa: A visual question answering benchmark requiring external knowledge. In: IEEE Conf. Comput. Vis. Pattern Recog. (2019)
2019
Earlier work this paper cites.
Reimers, N., Gurevych, I.: Sentence-bert: Sentence embeddings using siamese bert-networks. In: EMNLP (2019)
2019
Earlier work this paper cites.
Tan, H., Bansal, M.: Lxmert: Learning cross-modality encoder representations from transformers. In: EMNLP (2019)
2019
Earlier work this paper cites.
Zellers, R., Bisk, Y., Farhadi, A., Choi, Y.: From recognition to cognition: Visual commonsense reasoning. In: IEEE Conf. Comput. Vis. Pattern Recog. (2019)
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners. In: Adv. Neural Inform. Process. Syst. (2020)
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
Gardères, F., Ziaeefard, M., Abeloos, B., Lecue, F.: ConceptBert: Concept-aware representation for visual question answering. In: EMNLP (2020)
2020
Cited alongside, same era.
Selvaraju, R.R., Tendulkar, P., Parikh, D., Horvitz, E., Ribeiro, M.T., Nushi, B., Kamar, E.: Squinting at vqa models: Introspecting vqa models with sub-questions. In: IEEE Conf. Comput. Vis. Pattern Recog. (2020)
2020
Cited alongside, same era.
Zhu, Z., Yu, J., Wang, Y., Sun, Y., Hu, Y., Wu, Q.: Mucko: multi-layer cross-modal knowledge reasoning for fact-based visual question answering. In: IJCAI (2020)
2020
Cited alongside, same era.
Shen, S., Li, L.H., Tan, H., Bansal, M., Rohrbach, A., Chang, K.W., Yao, Z., Keutzer, K.: How much can clip benefit vision-and-language tasks? In: Int. Conf. Learn. Represent. (2022)
2022
Later among the works it cites.
Uehara, K., Duan, N., Harada, T.: Learning to ask informative sub-questions for visual question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. (2022)
2022
Later among the works it cites.
Yang, Z., Gan, Z., Wang, J., Hu, X., Lu, Y., Liu, Z., Wang, L.: An empirical study of gpt-3 for few-shot knowledge-based vqa. In: AAAI (2022)
2022
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Changpinyo, S., Sharma, P., Ding, N., Soricut, R.: Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In: IEEE Conf. Comput. Vis. Pattern Recog. (2021)
2021
Cited alongside, same era.
Dancette, C., Cadene, R., Teney, D., Cord, M.: Beyond question-based biases: Assessing multimodal shortcut learning in visual question answering. In: Int. Conf. Comput. Vis. (2021)
2021
Cited alongside, same era.
Marino, K., Chen, X., Parikh, D., Gupta, A., Rohrbach, M.: Krisp: Integrating implicit and symbolic knowledge for open-domain knowledge-based vqa. In: IEEE Conf. Comput. Vis. Pattern Recog. (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: Int. Conf. Mach. Learn. (2021)
2021
Cited alongside, same era.
Wu, J., Lu, J., Sabharwal, A., Mottaghi, R.: Multi-modal answer validation for knowledge-based vqa. In: AAAI (2021)
2021
Cited alongside, same era.
Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al.: Flamingo: a visual language model for few-shot learning. In: Adv. Neural Inform. Process. Syst. (2022)
2022
Cited alongside, same era.
Gao, F., Ping, Q., Thattai, G., Reganti, A., Wu, Y.N., Natarajan, P.: Transform-retrieve-generate: Natural language-centric outside-knowledge visual question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. (2022)
2022
Cited alongside, same era.
Guo, J., Li, J., Li, D., Tiong, A.M.H., Li, B., Tao, D., Hoi, S.C.: From images to textual prompts: Zero-shot vqa with frozen large language models. In: IEEE Conf. Comput. Vis. Pattern Recog. (2023)
2023
Later among the works it cites.
Hu, Z., Iscen, A., Sun, C., Wang, Z., Chang, K.W., Sun, Y., Schmid, C., Ross, D.A., Fathi, A.: Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge memory. In: IEEE Conf. Comput. Vis. Pattern Recog. (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Khan, Z., BG, V.K., Schulter, S., Yu, X., Fu, Y., Chandraker, M.: Q: How to specialize large vision-language models to data-scarce vqa tasks? a: Self-train on unlabeled images! In: IEEE Conf. Comput. Vis. Pattern Recog. (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Li, J., Li, D., Savarese, S., Hoi, S.: Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In: Int. Conf. Mach. Learn. (2023)
2023
Later among the works it cites.
Lin, W., Chen, J., Mei, J., Coca, A., Byrne, B.: Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering. In: Adv. Neural Inform. Process. Syst. (2023)
2023
Later among the works it cites.
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. In: Adv. Neural Inform. Process. Syst. (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Merullo, J., Castricato, L., Eickhoff, C., Pavlick, E.: Linearly mapping from image to text space. In: Int. Conf. Learn. Represent. (2023)
2023
Later among the works it cites.
Shao, Z., Yu, Z., Wang, M., Yu, J.: Prompting large language models with answer heuristics for knowledge-based visual question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.