Fetching the paper…
Reading the bibliography…
A key aspect of VQA models that are interpretable is their ability to ground their answers to relevant regions in the image.
The proof and measurement of association between two things
C. Spearman · 1904
Earlier work this paper cites.
Wordnet: a lexical database for english
G. A. Miller · 1995
Earlier work this paper cites.
Nltk: the natural language toolkit
S. Bird and E. Loper · 2004
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Visual turing test for computer vision systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Earlier work this paper cites.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
A. Das, H. Agrawal, C. L. Zitnick, D. Parikh, and D. Batra · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Hadamard product for low-rank bilinear pooling
J. Kim, K. W. On, W. Lim, J. Kim, J. Ha, and B. Zhang · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Image question answering using convolutional neural network with dynamic parameter prediction
H. Noh, P. Hongsuck Seo, and B. Han · 2016
Cited alongside, same era.
Where to look: Focus regions for visual question answering
K. J. Shih, S. Singh, and D. Hoiem · 2016
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Later among the works it cites.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
D. Teney, P. Anderson, X. He, and A. v. d. Hengel · 2017
Later among the works it cites.
Visual question answering: A survey of methods and datasets
Q. Wu, D. Teney, P. Wang, C. Shen, A. Dick, and A. van den Hengel · 2017
Later among the works it cites.
Multi-modal factorized bilinear pooling with co-attention learning for visual question answering
Z. Yu, J. Yu, J. Fan, and D. Tao · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Xu and K. Saenko · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and VQA
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2017
Cited alongside, same era.
Vqs: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation
C. Gan, Y. Li, H. Li, C. Sun, and B. Gong · 2017
Cited alongside, same era.
Beyond bilinear: Generalized multi-modal factorized high-order pooling for visual question answering
Z. Yu, J. Yu, C. Xiang, J. Fan, and D. Tao · 2017
Later among the works it cites.
Multimodal explanations: Justifying decisions and pointing to the evidence
D. H. Park, L. A. Hendricks, Z. Akata, A. Rohrbach, B. Schiele, T. Darrell, and M. Rohrbach · 2018
Closest in time.
Exploring human-like attention supervision in visual question answering
T. Qiao, J. Dong, and D. Xu · 2018
Closest in time.