Fetching the paper…
Reading the bibliography…
Visual question answering (VQA) requires joint comprehension of images and natural language questions, where many questions can't be directly or clearly answered from visual content but require reasoning from structured human knowledge with confirmation from visual content.
Building a question answering test collection
E. M. Voorhees and D. M. Tice · 2000
Earlier work this paper cites.
Conceptnet – a practical commonsense reasoning toolkit
H. Liu and P. Singh · 2004
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
S. Auer, C. Bizer, G. Kobilarov, et al · 2007
Earlier work this paper cites.
Freebase: a collaboratively created graph database for structuring human knowledge
K. Bollacker, C. Evans, P. Paritosh, et al · 2008
Earlier work this paper cites.
A survey on question answering technology from an information retrieval perspective
O. Kolomiyets and M.-F. Moens · 2011
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
J. Berant, A. Chou, R. Frostig, and P. Liang · 2013
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
A. Bordes, N. Usunier, A. Garcia-Duran, J. Weston, and O. Yakhnenko · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Question answering with subgraph embeddings
A. Bordes, S. Chopra, and J. Weston · 2014
Earlier work this paper cites.
A Fast and Accurate Dependency Parser using Neural Networks
D. Chen and C. D. Manning · 2014
Earlier work this paper cites.
Open question answering over curated and extracted knowledge bases
A. Fader, L. Zettlemoyer, and O. Etzioni · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T. Y. Lin, M. Maire, S. Belongie, J. Hays, et al · 2014
Earlier work this paper cites.
Towards a visual turing challenge
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
J. Weston, S. Chopra, and A. Bordes · 2014
Earlier work this paper cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, et al · 2015
Earlier work this paper cites.
Visual turing test for computer vision systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Earlier work this paper cites.
RNN : Recurrent Library for Torch
N. Léonard, S. Waghmare, Y. Wang, and J.-H. Kim · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Cited alongside, same era.
End-to-end memory networks
S. Sukhbaatar, J. Weston, R. Fergus, et al · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, et al · 2015
Cited alongside, same era.
Abcnn: Attention-based convolutional neural network for modeling sentence pairs
W. Yin, H. Schütze, B. Xiang, and B. Zhou · 2015
Cited alongside, same era.
Simple Baseline for Visual Question Answering
B. Zhou, Y. Tian, S. Sukhbaatar, A. Szlam, and R. Fergus · 2015
Cited alongside, same era.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Q. Wu, P. Wang, C. Shen, A. Dick, et al · 2016
Later among the works it cites.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Later among the works it cites.
http://www.visualqa.org/roe_2017.html
Vqa 2.0 challenge leaderboard, 2017 · 2017
Later among the works it cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, et al · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Zhu et al · 2015
Cited alongside, same era.
VQA: Visual Question Answering
A. Agrawal, J. Lu, S. Antol, M. Mitchell, L. Zitnick, et al · 2016
Cited alongside, same era.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
C. N. dos Santos, M. Tan, B. Xiang, and B. Zhou · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, et al · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Multimodal Residual Learning for Visual QA
J.-H. Kim, S.-W. Lee, D.-H. Kwak, et al · 2016
Cited alongside, same era.
Later among the works it cites.
Snapshot ensembles: Train 1, get m for free
G. Huang, Y. Li, G. Pleiss, Z. Liu, et al · 2017
Later among the works it cites.
Hadamard product for low-rank bilinear pooling
J.-H. Kim, K.-W. On, J. Kim, J.-W. Ha, and B.-T. Zhang · 2017
Later among the works it cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, et al · 2017
Later among the works it cites.
Dual attention networks for multimodal reasoning and matching
H. Nam, J.-W. Ha, and J. Kim · 2017
Later among the works it cites.
Weakly supervised dense video captioning
Z. Shen, J. Li, Z. Su, M. Li, Y. Chen, Y.-G. Jiang, and X. Xue · 2017
Later among the works it cites.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
D. Teney, P. Anderson, X. He, and A. v. d. Hengel · 2017
Later among the works it cites.
Explicit knowledge-based reasoning for visual question answering
P. Wang, Q. Wu, et al · 2017
Later among the works it cites.
Fvqa: Fact-based visual question answering
P. Wang, Q. Wu, C. Shen, et al · 2017
Later among the works it cites.
Multi-modal factorized bilinear pooling with co-attention learning for visual question answering
Z. Yu, J. Yu, J. Fan, and D. Tao · 2017
Later among the works it cites.
Structured attentions for visual question answering
C. Zhu, Y. Zhao, S. Huang, K. Tu, and Y. Ma · 2017
Later among the works it cites.
Cascade r-cnn: Delving into high quality object detection
Z. Cai and N. Vasconcelos · 2018
Closest in time.