Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) deep-learning systems tend to capture superficial statistical correlations in the training data because of strong language priors and fail to generalize to test data with a significantly different question-answer (QA) distribution.
Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Glove: Global Vectors for Word Representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-time Object Detection with Region Proposal Networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Neural Module Networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Stacked Attention Networks for Image Question Answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Earlier work this paper cites.
MUTAN: Multimodal Tucker Fusion for Visual Question Answering
H. Ben-Younes, R. Cadene, M. Cord, and N. Thome · 2017
Earlier work this paper cites.
Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?
A. Das, H. Agrawal, L. Zitnick, D. Parikh, and D. Batra · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Cited alongside, same era.
spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing
M. Honnibal and I. Montani · 2017
Cited alongside, same era.
Right for the Right Reasons: Training Differentiable Models by Constraining Their Explanations
A. S. Ross, M. C. Hughes, and F. Doshi-Velez · 2017
Cited alongside, same era.
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, et al · 2017
Cited alongside, same era.
Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge
D. Teney, P. Anderson, X. He, and A. v. d. Hengel · 2017
Cited alongside, same era.
Exploring Human-Like Attention Supervision in Visual Question Answering
T. Qiao, J. Dong, and D. Xu · 2018
Later among the works it cites.
Overcoming Language Priors in Visual Question Answering with Adversarial Regularization
S. Ramakrishnan, A. Agrawal, and S. Lee · 2018
Later among the works it cites.
Interpretable Counting for Visual Question Answering
A. Trott, C. Xiong, and R. Socher · 2018
Later among the works it cites.
Dynamic Filtering with Large Sampling Field for Convnets
J. Wu, D. Li, Y. Yang, C. Bajaj, and X. Ji · 2018
Later among the works it cites.
Attention is not Explanation
S. Jain and B. C. Wallace · 2019
Closest in time.
Taking a HINT: Leveraging Explanations to Make Vision and Language Models More Grounded
R. R. Selvaraju, S. Lee, Y. Shen, H. Jin, D. Batra, and D. Parikh · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi · 2018
Cited alongside, same era.
Bottom-Up and Top-Down Attention for Image Captioning and VQA
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Cited alongside, same era.
Explainable Neural Computation via Stack Neural Module Networks
R. Hu, J. Andreas, T. Darrell, and K. Saenko · 2018
Cited alongside, same era.
Pythia v0. 1: the Winning Entry to the VQA Challenge 2018
Y. Jiang, V. Natarajan, X. Chen, M. Rohrbach, D. Batra, and D. Parikh · 2018
Cited alongside, same era.
Bilinear Attention Networks
J.-H. Kim, J. Jun, and B.-T. Zhang · 2018
Cited alongside, same era.
Multimodal Explanations: Justifying Decisions and Pointing to the Evidence
D. H. Park, L. A. Hendricks, Z. Akata, A. Rohrbach, B. Schiele, T. Darrell, and M. Rohrbach · 2018
Cited alongside, same era.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al
Cited in the paper.
Cycle-Consistency for Robust Visual Question Answering
M. Shah, X. Chen, M. Rohrbach, and D. Parikh · 2019
Closest in time.
Generating Question Relevant Captions to Aid Visual Question Answering
J. Wu, Z. Hu, and R. J. Mooney · 2019
Closest in time.
Faithful Multimodal Explanation for Visual Question Answering
J. Wu and R. J. Mooney · 2019
Closest in time.
Visual Entailment: A Novel Task for Fine-Grained Image Understanding
N. Xie, F. Lai, D. Doran, and A. Kadav · 2019
Closest in time.
Interpretable Visual Question Answering by Visual Grounding from Attention Supervision Mining
Y. Zhang, J. C. Niebles, and A. Soto · 2019
Closest in time.