Fetching the paper…
Reading the bibliography…
Visual question answering requires high-order reasoning about an image, which is a fundamental capability needed by machine systems to follow complex directives.
Matplotlib: A 2d graphics environment
J. D. Hunter · 2007
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Deep compositional question answering with neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2015
Earlier work this paper cites.
ABC-CNN: an attention based convolutional neural network for visual question answering
K. Chen, J. Wang, L. Chen, H. Gao, W. Xu, and R. Nevatia · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Where To Look: Focus Regions for Visual Question Answering
K. J. Shih, S. Singh, and D. Hoiem · 2015
Earlier work this paper cites.
Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering
H. Xu and K. Saenko · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Earlier work this paper cites.
Stacked Attention Networks for Image Question Answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2015
Earlier work this paper cites.
Visual7W: Grounded Question Answering in Images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2015
Earlier work this paper cites.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
A. Fukui, D. Huk Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. B. Girshick · 2016
Cited alongside, same era.
The Mythos of Model Interpretability
Z. C. Lipton · 2016
Cited alongside, same era.
Hierarchical Question-Image Co-Attention for Visual Question Answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Inferring and executing programs for visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, J. Hoffman, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Later among the works it cites.
N. Liu and J. Han · 2017
Later among the works it cites.
Learning Visual Reasoning Without Strong Priors
E. Perez, H. de Vries, F. Strub, V. Dumoulin, and A. Courville · 2017
Later among the works it cites.
Person Re-identification Using Visual Attention
A. Rahimpour, L. Liu, A. Taalimi, Y. Song, and H. Qi · 2017
Later among the works it cites.
A simple neural network module for relational reasoning
A. Santoro, D. Raposo, D. G. T. Barrett, M. Malinowski, R. Pascanu, P. Battaglia, and T. Lillicrap · 2017
Later among the works it cites.
Residual Attention Network for Image Classification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra · 2016
Cited alongside, same era.
Multi-scale context aggregation by dilated convolutions
F. Yu and V. Koltun · 2016
Cited alongside, same era.
Bottom-Up and Top-Down Attention for Image Captioning and VQA
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2017
Cited alongside, same era.
Segmentation-aware convolutional networks using local attention masks
A. W. Harley, K. G. Derpanis, and I. Kokkinos · 2017
Cited alongside, same era.
Learning to reason: End-to-end module networks for visual question answering
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko · 2017
Cited alongside, same era.
Video summarization with attention-based encoder-decoder networks
Z. Ji, K. Xiong, Y. Pang, and X. Li · 2017
Cited alongside, same era.
F. Wang, M. Jiang, C. Qian, S. Yang, C. Li, H. Zhang, X. Wang, and X. Tang · 2017
Later among the works it cites.
Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
Z. Yu, J. Yu, J. Fan, and D. Tao · 2017
Later among the works it cites.
Structured Attentions for Visual Question Answering
C. Zhu, Y. Zhao, S. Huang, K. Tu, and Y. Ma · 2017
Later among the works it cites.
Compositional attention networks for machine reasoning
C. D. M. Drew Arad Hudson · 2018
Closest in time.
DDRprog: A CLEVR differentiable dynamic reasoning programmer, 2018
J. Suarez, J. Johnson, and F.-F. Li · 2018
Closest in time.
Learning to count objects in natural images for visual question answering
Y. Zhang, J. Hare, and A. Prügel-Bennett · 2018
Closest in time.