Fetching the paper…
Reading the bibliography…
A number of studies have found that today's Visual Question Answering (VQA) models are heavily driven by superficial correlations in the training data and lack sufficient image grounding.
Relevance feedback in image retrieval: A comprehensive review
X. S. Zhou and T. S. Huang · 2003
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer
C. H. Lampert, H. Nickisch, and S. Harmeling · 2009
Earlier work this paper cites.
Torch7: A matlab-like environment for machine learning
R. Collobert, K. Kavukcuoglu, and C. Farabet · 2011
Earlier work this paper cites.
Decorrelating semantic visual attributes by resisting the urge to share
D. Jayaraman, F. Sha, and K. Grauman · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
ABC-CNN: an attention based convolutional neural network for visual question answering
K. Chen, J. Wang, L. Chen, H. Gao, W. Xu, and R. Nevatia · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Earlier work this paper cites.
Visual turing test for computer vision systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Earlier work this paper cites.
Compositional memory for visual question answering
A. Jiang, F. Wang, F. Porikli, and Y. Li · 2015
Earlier work this paper cites.
Deeper lstm and normalized cnn visual question answering model
J. Lu, X. Lin, D. Batra, and D. Parikh · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Earlier work this paper cites.
Explicit knowledge-based reasoning for visual question answering
P. Wang, Q. Wu, C. Shen, A. van den Hengel, and A. R. Dick · 2015
Cited alongside, same era.
Analyzing the behavior of visual question answering models
A. Agrawal, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Deep compositional question answering with neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Learning to generalize to new compositions in image understanding
Y. Atzmon, J. Berant, V. Kezami, A. Globerson, and G. Chechik · 2016
Cited alongside, same era.
Zero-shot visual question answering
D. Teney and A. v. d. Hengel · 2016
Later among the works it cites.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Q. Wu, P. Wang, C. Shen, A. van den Hengel, and A. R. Dick · 2016
Later among the works it cites.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. J. Smola · 2016
Later among the works it cites.
Yin and Yang: Balancing and answering binary visual questions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
A focused dynamic attention model for visual question answering
I. Ilievski, S. Yan, and J. Feng · 2016
Cited alongside, same era.
Answer-type prediction for visual question answering
K. Kafle and C. Kanan · 2016
Cited alongside, same era.
Multimodal residual learning for visual QA
J.-H. Kim, S.-W. Lee, D.-H. Kwak, M.-O. Heo, J. Kim, J.-W. Ha, and B.-T. Zhang · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Training recurrent answering units with joint loss minimization for vqa
H. Noh and B. Han · 2016
Cited alongside, same era.
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Later among the works it cites.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
C-vqa: A compositional split of the visual question answering (vqa) v1. 0 dataset
A. Agrawal, A. Kembhavi, D. Batra, and D. Parikh · 2017
Closest in time.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Closest in time.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Closest in time.
An analysis of visual question answering algorithms
K. Kafle and K. Christopher · 2017
Closest in time.
An empirical evaluation of visual question answering for novel objects
S. K. Ramakrishnan, A. Pal, G. Sharma, and A. Mittal · 2017
Closest in time.
Being negative but constructively: Lessons learnt from creating better visual question answering datasets
W. Chao, H. Hu, and F. Sha · 2018
Closest in time.
Cross-dataset adaptation for visual question answering
W. Chao, H. Hu, and F. Sha · 2018
Closest in time.