Fetching the paper…
Reading the bibliography…
Knowledge-based visual question answering (KVQA) has been extensively studied to answer visual questions with external knowledge, e.g., knowledge graphs (KGs).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Pseudo siamese network for few-shot intent generation
Congying Xia, Caiming Xiong, and Philip Yu. 2021 · 2009
Earlier work this paper cites.
Mercer’s theorem on general domains: On the interaction between measures, kernels, and rkhss
Ingo Steinwart and Clint Scovel. 2012 · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
Denny Vrandečić and Markus Krötzsch. 2014 · 2014
Earlier work this paper cites.
Mutan: Multimodal tucker fusion for visual question answering
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome. 2017 · 2017
Earlier work this paper cites.
Fvqa: Fact-based visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. 2017 · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. 2018 · 2018
Earlier work this paper cites.
Bilinear attention networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. 2018 · 2018
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi. 2018 · 2018
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019 · 2019
Earlier work this paper cites.
Explainable and explicit visual reasoning over scene graphs
Jiaxin Shi, Hanwang Zhang, and Juanzi Li. 2019 · 2019
Earlier work this paper cites.
Deep modular co-attention networks for visual question answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. 2019 · 2019
Cited alongside, same era.
Conceptbert: Concept-aware representation for visual question answering
François Gardères, Maryam Ziaeefard, Baptiste Abeloos, and Freddy Lecue. 2020 · 2020
Cited alongside, same era.
Cross-modal knowledge reasoning for knowledge-based visual question answering
Jing Yu, Zihao Zhu, Yujing Wang, Weifeng Zhang, Yue Hu, and Jianlong Tan. 2020 · 2020
Cited alongside, same era.
Krisp: Integrating implicit and symbolic knowledge for open-domain knowledge-based vqa
Kenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta, and Marcus Rohrbach. 2021 · 2021
Cited alongside, same era.
Explicit knowledge incorporation for visual reasoning
Yifeng Zhang, Ming Jiang, and Qi Zhao. 2021 · 2021
Cited alongside, same era.
Neighbor enhanced graph convolutional networks for node classification and recommendation
Hierarchy-aware multi-hop question answering over knowledge graphs
Junnan Dong, Qinggang Zhang, Xiao Huang, Keyu Duan, Qiaoyu Tan, and Zhimeng Jiang. 2023 · 2023
Later among the works it cites.
Learning to fake it: limited responses and fabricated references provided by chatgpt for medical questions
Jocelyn Gravel, Madeleine D’Amours-Gravel, and Esli Osmanlliu. 2023 · 2023
Later among the works it cites.
Siamese masked autoencoders
Agrim Gupta, Jiajun Wu, Jia Deng, and Li Fei-Fei. 2023 · 2023
Later among the works it cites.
Promptcap: Prompt-guided task-aware image captioning
Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi, Noah A Smith, and Jiebo Luo. 2023 · 2023
Later among the works it cites.
Vlc-bert: visual question answering with contextualized commonsense knowledge
Sahithya Ravi, Aditya Chinchure, Leonid Sigal, Renjie Liao, and Vered Shwartz. 2023 · 2023
Later among the works it cites.
In chatgpt we trust? measuring and characterizing the reliability of chatgpt
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hao Chen, Zhong Huang, Yue Xu, Zengde Deng, Feiran Huang, Peng He, and Zhoujun Li. 2022 · 2022
Cited alongside, same era.
Kat: A knowledge augmented transformer for vision-and-language
Liangke Gui, Borui Wang, Qiuyuan Huang, Alexander G Hauptmann, Yonatan Bisk, and Jianfeng Gao. 2022 · 2022
Cited alongside, same era.
G-mixup: Graph data augmentation for graph classification
Xiaotian Han, Zhimeng Jiang, Ninghao Liu, and Xia Hu. 2022 · 2022
Cited alongside, same era.
Revive: Regional visual representation matters in knowledge-based visual question answering
Yuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu, Chenguang Zhu, and Lu Yuan. 2022 · 2022
Cited alongside, same era.
Multi-modal answer validation for knowledge-based vqa
Jialin Wu, Jiasen Lu, Ashish Sabharwal, and Roozbeh Mottaghi. 2022 · 2022
Cited alongside, same era.
An empirical study of gpt-3 for few-shot knowledge-based vqa
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang. 2022 · 2022
Cited alongside, same era.
Ai unreliable answers: A case study on chatgpt
Ilaria Amaro, Attilio Della Greca, Rita Francese, Genoveffa Tortora, and Cesare Tucci. 2023 · 2023
Cited alongside, same era.
Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023 · 2023
Later among the works it cites.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. 2023 · 2023
Later among the works it cites.
Toward multi-granularity decision-making: Explicit visual reasoning with hierarchical knowledge
Yifeng Zhang, Shi Chen, and Qi Zhao. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Later among the works it cites.
Knowledge graphs meet multi-modal learning: A comprehensive survey
Zhuo Chen, Yichi Zhang, Yin Fang, Yuxia Geng, Lingbing Guo, Xiang Chen, Qian Li, Wen Zhang, Jiaoyan Chen, Yushan Zhu, et al. 2024 · 2024
Closest in time.
Knowledge-to-sql: Enhancing sql generation with data expert llm
Zijin Hong, Zheng Yuan, Hao Chen, Qinggang Zhang, Feiran Huang, and Xiao Huang. 2024 · 2024
Closest in time.
Differentiable neuro-symbolic reasoning on large-scale knowledge graphs
Chen Shengyuan, Yunfeng Cai, Huang Fang, Xiao Huang, and Mingming Sun. 2024 · 2024
Closest in time.
A unified end-to-end retriever-reader framework for knowledge-based vqa
Yangyang Guo, Liqiang Nie, Yongkang Wong, Yibing Liu, Zhiyong Cheng, and Mohan Kankanhalli. 2022 · 2069
Closest in time.