Fetching the paper…
Reading the bibliography…
Knowledge-based visual question answering (VQA) is a vision-language task that requires an agent to correctly answer image-related questions using knowledge that is not presented in the given image.
GQA: a new dataset for compositional question answering over real-world images
Hudson, D. A.; and Manning, C. D. 2019 · 1902
Earlier work this paper cites.
Explainable high-order visual question reasoning: A new benchmark and knowledge-routed network
Cao, Q.; Li, B.; Liang, X.; and Lin, L. 2019 · 1909
Earlier work this paper cites.
Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks
Wang, P.; Wu, Q.; Cao, J.; Shen, C.; Gao, L.; and van den Hengel, A. 2019 · 1968
Earlier work this paper cites.
Scene graph generation with external knowledge and image reconstruction
Gu, J.; Zhao, H.; Lin, Z.; Li, S.; Cai, J.; and Ling, M. 2019 · 1978
Earlier work this paper cites.
DBpedia: A nucleus for a web of open data
Auer, S.; Bizer, C.; Kobilarov, G.; Lehmann, J.; Cyganiak, R.; and Ives, Z. 2007 · 2007
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Earlier work this paper cites.
Webchild: Harvesting and organizing commonsense knowledge from the Web
Tandon, N.; De Melo, G.; Suchanek, F.; and Weikum, G. 2014 · 2014
Earlier work this paper cites.
Weston, J.; Chopra, S.; and Bordes, A. 2014 · 2014
Earlier work this paper cites.
VQA: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Sukhbaatar, S.; Szlam, A.; Weston, J.; and Fergus, R. 2015 · 2015
Earlier work this paper cites.
Neural module networks
Andreas, J.; Rohrbach, M.; Darrell, T.; and Klein, D. 2016 · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A.; Park, D. H.; Yang, D.; Rohrbach, A.; Darrell, T.; and Rohrbach, M. 2016 · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Lu, J.; Yang, J.; Batra, D.; and Parikh, D. 2016 · 2016
Cited alongside, same era.
Learning to answer questions from image using convolutional neural network
Ma, L.; Lu, Z.; and Li, H. 2016 · 2016
Cited alongside, same era.
Key-value memory networks for directly reading documents
Miller, A.; Fisch, A.; Dodge, J.; Karimi, A.-H.; Bordes, A.; and Weston, J. 2016 · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Yang, Z.; He, X.; Gao, J.; Deng, L.; and Smola, A. 2016 · 2016
Cited alongside, same era.
MUTAN: Multimodal tucker fusion for visual question answering
Ben-younes, H.; Cadene, R.; Cord, M.; and Thome, N. 2017 · 2017
Cited alongside, same era.
Conceptnet 5.5: An open multilingual graph of general knowledge
Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering
Yu, Z.; Yu, J.; Xiang, C.; Fan, J.; and Tao, D. 2018 · 2018
Later among the works it cites.
Language-conditioned graph networks for relational reasoning
Hu, R.; Rohrbach, A.; Darrell, T.; and Saenko, K. 2019 · 2019
Later among the works it cites.
OK-VQA: A visual question answering benchmark requiring external knowledge
Marino, K.; Rastegari, M.; Farhadi, A.; and Mottaghi, R. 2019 · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Later among the works it cites.
Enhancing key-value memory neural networks for knowledge based question answering
Xu, K.; Lai, Y.; Feng, Y.; and Wang, Z. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Speer, R.; Chin, J.; and Havasi, C. 2017 · 2017
Cited alongside, same era.
Graph-structured representations for visual question answering
Teney, D.; Liu, L.; and van den Hengel, A. 2017 · 2017
Cited alongside, same era.
Graph attention networks
Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Liò, P.; and Bengio, Y. 2017 · 2017
Cited alongside, same era.
FVQA: Fact-based visual question answering
Wang, P.; Wu, Q.; Shen, C.; Dick, A.; and van den Hengel, A. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Cited alongside, same era.
Out of the box: Reasoning with graph convolution nets for factual visual question answering
Narasimhan, M.; Lazebnik, S.; and Schwing, A. G. 2018 · 2018
Cited alongside, same era.
Straight to the facts: Learning knowledge base retrieval for factual visual question answering
Narasimhan, M.; and Schwing, A. G. 2018 · 2018
Cited alongside, same era.
Yu, Z.; Yu, J.; Cui, Y.; Tao, D.; and Tian, Q. 2019 · 2019
Later among the works it cites.
Stanza: A Python natural language processing toolkit for many human languages
Qi, P.; Zhang, Y.; Zhang, Y.; Bolton, J.; and Manning, C. D. 2020 · 2020
Later among the works it cites.
Cross-modal knowledge reasoning for knowledge-based visual question answering
Yu, J.; Zhu, Z.; Wang, Y.; Zhang, W.; Hu, Y.; and Tan, J. 2020 · 2020
Later among the works it cites.
Bridging knowledge graphs to generate scene graphs
Zareian, A.; Karaman, S.; and Chang, S.-F. 2020 · 2020
Later among the works it cites.
Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering
Zhu, Z.; Yu, J.; Wang, Y.; Sun, Y.; Hu, Y.; and Wu, Q. 2020 · 2020
Later among the works it cites.
Towards knowledge-augmented visual question answering
Ziaeefard, M.; and Lécué, F. 2020 · 2020
Later among the works it cites.
Knowledge-routed visual question reasoning: Challenges for deep representation embedding
Cao, Q.; Li, B.; Liang, X.; Wang, K.; and Lin, L. 2021 · 2021
Later among the works it cites.
Select, substitute, search: A new benchmark for knowledge-augmented visual question answering
Jain, A.; Kothyari, M.; Kumar, V.; Jyothi, P.; Ramakrishnan, G.; and Chakrabarti, S. 2021 · 2021
Later among the works it cites.