Fetching the paper…
Reading the bibliography…
The problem of knowledge-based visual question answering involves answering questions that require external knowledge in addition to the content of the image.
Well-read students learn better: On the importance of pre-training compact models
Turc, I.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 1908
Earlier work this paper cites.
ReferItGame: Referring to Objects in Photographs of Natural Scenes
Kazemzadeh, S.; Ordonez, V.; Matten, M.; and Berg, T. 2014 · 2014
Earlier work this paper cites.
Joint training of a convolutional network and a graphical model for human pose estimation
Tompson, J. J.; Jain, A.; LeCun, Y.; and Bregler, C. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Hierarchical Question-Image Co-attention for Visual Question Answering
Lu, J.; Yang, J.; Batra, D.; and Parikh, D. 2016 · 2016
Earlier work this paper cites.
Ask Me Anything: Free-Form Visual Question Answering Based on Knowledge From External Sources
Wu, Q.; Wang, P.; Shen, C.; Dick, A.; and van den Hengel, A. 2016 · 2016
Earlier work this paper cites.
Visual7w: Grounded Question Answering in Images
Zhu, Y.; Groth, O.; Bernstein, M.; and Fei-Fei, L. 2016 · 2016
Earlier work this paper cites.
MUTAN: Multimodal Tucker Fusion for Visual Question Answering
Ben-Younes, H.; Cadene, R.; Cord, M.; and Thome, N. 2017 · 2017
Earlier work this paper cites.
Mask r-cnn
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. 2017 · 2017
Earlier work this paper cites.
Answering Complex Questions Using Open Information Extraction
Khot, T.; Sabharwal, A.; and Clark, P. 2017 · 2017
Earlier work this paper cites.
End-to-end neural coreference resolution
Lee, K.; He, L.; Lewis, M.; and Zettlemoyer, L. 2017 · 2017
Earlier work this paper cites.
Explicit knowledge-based reasoning for visual question answering
Wang, P.; Wu, Q.; Shen, C.; Hengel, A. v. d.; and Dick, A. 2017 · 2017
Earlier work this paper cites.
Bottom-Up and Top-Down Attention for Image Captioning and VQA
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Earlier work this paper cites.
Transforming question answering datasets into natural language inference datasets
Demszky, D.; Guu, K.; and Liang, P. 2018 · 2018
Earlier work this paper cites.
Bilinear Attention Networks
Kim, J.-H.; Jun, J.; and Zhang, B.-T. 2018 · 2018
Earlier work this paper cites.
Out-of-The-Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering
Narasimhan, M.; Lazebnik, S.; and Schwing, A. 2018 · 2018
Cited alongside, same era.
Straight to the facts: Learning knowledge base retrieval for factual visual question answering
Narasimhan, M.; and Schwing, A. G. 2018 · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Cited alongside, same era.
Fvqa: Fact-based visual question answering
Wang, P.; Wu, Q.; Shen, C.; Dick, A.; and van den Hengel, A. 2018 · 2018
Cited alongside, same era.
Murel: Multimodal relational reasoning for visual question answering
Cadene, R.; Ben-Younes, H.; Cord, M.; and Thome, N. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
ConceptBert: Concept-Aware Representation for Visual Question Answering
Gardères, F.; Ziaeefard, M.; Abeloos, B.; and Lecue, F. 2020 · 2020
Later among the works it cites.
Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-Training
Li, G.; Duan, N.; Fang, Y.; Gong, M.; Jiang, D.; and Zhou, M. 2020 · 2020
Later among the works it cites.
Boosting Visual Question Answering with Context-aware Knowledge Aggregation
Li, G.; Wang, X.; and Zhu, W. 2020 · 2020
Later among the works it cites.
12-in-1: Multi-Task Vision and Language Representation Learning
Lu, J.; Goswami, V.; Rohrbach, M.; Parikh, D.; and Lee, S. 2020 · 2020
Later among the works it cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
GQA: a new dataset for compositional question answering over real-world images
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Cited alongside, same era.
Progressive Attention Memory Network for Movie Story Question Answering
Kim, J.; Ma, M.; Kim, K.; Kim, S.; and Yoo, C. D. 2019 · 2019
Cited alongside, same era.
VisualBERT: A simple and performant baseline for vision and language
Li, L. H.; Yatskar, M.; Yin, D.; Hsieh, C.-J.; and Chang, K.-W. 2019 · 2019
Cited alongside, same era.
Learning Rich Image Region Representation for Visual Question Answering
Liu, B.; Huang, Z.; Zeng, Z.; Chen, Z.; and Fu, J. 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Cited alongside, same era.
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
Marino, K.; Rastegari, M.; Farhadi, A.; and Mottaghi, R. 2019 · 2019
Cited alongside, same era.
Wu, J.; Chen, L.; and Mooney, R. J. 2020 · 2020
Later among the works it cites.
BERTScore: Evaluating Text Generation with BERT
Zhang*, T.; Kishore*, V.; Wu*, F.; Weinberger, K. Q.; and Artzi, Y. 2020 · 2020
Later among the works it cites.
AnswerFact: Fact Checking in Product Question Answering
Zhang, W.; Deng, Y.; Ma, J.; and Lam, W. 2020 · 2020
Later among the works it cites.
Unified Vision-Language Pre-Training for Image Captioning and VQA
Zhou, L.; Palangi, H.; Zhang, L.; Hu, H.; Corso, J. J.; and Gao, J. 2020 · 2020
Later among the works it cites.
Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering
Zhu, Z.; Yu, J.; Wang, Y.; Sun, Y.; Hu, Y.; and Wu, Q. 2020 · 2020
Later among the works it cites.
Decontextualization: Making Sentences Stand-Alone
Choi, E.; Palomaki, J.; Lamm, M.; Kwiatkowski, T.; Das, D.; and Collins, M. 2021 · 2021
Closest in time.
KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQA
Marino, K.; Chen, X.; Parikh, D.; Gupta, A.; and Rohrbach, M. 2021 · 2021
Closest in time.
Passage Retrieval for Outside-Knowledge Visual Question Answering
Qu, C.; Zamani, H.; Yang, L.; Croft, W. B.; and Learned-Miller, E. 2021 · 2021
Closest in time.
Reasoning over Vision and Language: Exploring the Benefits of Supplemental Knowledge
Shevchenko, V.; Teney, D.; Dick, A.; and Hengel, A. v. d. 2021 · 2021
Closest in time.
Joint Models for Answer Verification in Question Answering Systems
Zhang, Z.; Vu, T.; and Moschitti, A. 2021 · 2021
Closest in time.