Fetching the paper…
Reading the bibliography…
When answering questions about an image, it not only needs knowing what -- understanding the fine-grained contents (e.g., objects, relationships) in the image, but also telling why -- reasoning over grounding visual cues to derive the answer for a question.
1908
Earlier work this paper cites.
1909
Earlier work this paper cites.
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler, “What are you talking about? text-to-image coreference,” in CVPR , 2014
2014
Earlier work this paper cites.
V. Ramanathan, A. Joulin, P. Liang, and L. Fei-Fei, “Linking people in videos with “their” names using coreference resolution,” in ECCV , 2014, pp. 95–110
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh, “Vqa: Visual question answering,” in ICCV , 2015
2015
Earlier work this paper cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in ICCV , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in NeurIPS , 2015
2015
Earlier work this paper cites.
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei, “Image retrieval using scene graphs,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3668–3678
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR , 2015
2015
Earlier work this paper cites.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” in NeurIPS , 2016
2016
Earlier work this paper cites.
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh, “Yin and yang: Balancing and answering binary visual questions,” in CVPR , 2016
2016
Earlier work this paper cites.
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei, “Visual7w: Grounded question answering in images,” in CVPR , 2016
2016
Earlier work this paper cites.
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein, “Neural module networks,” in CVPR , 2016
2016
Earlier work this paper cites.
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola, “Stacked attention networks for image question answering,” in CVPR , 2016
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in ECCV , 2016
2016
Earlier work this paper cites.
A. Jabri, A. Joulin, and L. van der Maaten, “Revisiting visual question answering baselines,” in ECCV , 2016
2016
Earlier work this paper cites.
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick, “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,” in CVPR , 2017
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in ICCV , 2017, pp. 618–626
2017
Earlier work this paper cites.
A. Das, H. Agrawal, L. Zitnick, D. Parikh, and D. Batra, “Human attention in visual question answering: Do humans and deep networks look at the same regions?” CVIU , vol. 163, pp. 90–100, 2017
2017
Earlier work this paper cites.
C. Gan, Y. Li, H. Li, C. Sun, and B. Gong, “Vqs: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation,” in ICCV , 2017
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in CVPR , 2017
2017
Earlier work this paper cites.
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko, “Learning to reason: End-to-end module networks for visual question answering,” in ICCV , 2017
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
J. Lu, C. Xiong, D. Parikh, and R. Socher, “Knowing when to look: Adaptive attention via a visual sentinel for image captioning,” in CVPR , 2017, pp. 375–383
2017
Cited alongside, same era.
D. Teney, L. Liu, and A. van Den Hengel, “Graph-structured representations for visual question answering,” in CVPR , 2017, pp. 1–9
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Cited alongside, same era.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in CVPR , 2017, pp. 1492–1500
B. N. Patro, M. Lunayach, S. Patel, and V. P. Namboodiri, “U-cam: Visual explanation using uncertainty based class activation maps,” in ICCV , 2019
2019
Later among the works it cites.
R. R. Selvaraju, S. Lee, Y. Shen, H. Jin, S. Ghosh, L. Heck, D. Batra, and D. Parikh, “Taking a hint: Leveraging explanations to make vision and language models more grounded,” in ICCV , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
R. Cadene, C. Dancette, M. Cord, D. Parikh et al. , “Rubi: Reducing unimodal biases for visual question answering,” in Advances in neural information processing systems , 2019, pp. 841–852
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi, “Don’t just assume; look and answer: Overcoming priors for visual question answering,” in CVPR , 2018
2018
Cited alongside, same era.
A. Singh, V. Goswami, V. Natarajan, Y. Jiang, X. Chen, M. Shah, M. Rohrbach, D. Batra, and D. Parikh, “Pythia-a platform for vision & language research,” in SysML Workshop, NeurIPS , vol. 2018, 2018
2018
Cited alongside, same era.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in CVPR , 2018
2018
Cited alongside, same era.
K. Yi, J. Wu, C. Gan, A. Torralba, P. Kohli, and J. B. Tenenbaum, “Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding,” in NeurIPS , 2018
2018
Cited alongside, same era.
D. A. Hudson and C. D. Manning, “Compositional attention networks for machine reasoning,” in ICLR , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,” in EMNLP , 2019
2019
Later among the works it cites.
J. Mao, C. Gan, P. Kohli, J. B. Tenenbaum, and J. Wu, “The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision,” ICLR , 2019
2019
Later among the works it cites.
Z. Fan, Z. Wei, S. Wang, and X.-J. Huang, “Bridging by word: Image grounded vocabulary construction for visual captioning,” in ACL , 2019
2019
Later among the works it cites.
R. Liu, C. Liu, Y. Bai, and A. L. Yuille, “Clevr-ref+: Diagnosing visual reasoning with referring expressions,” in CVPR , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
J.-H. Huang, C. D. Dao, M. Alfadly, and B. Ghanem, “A novel framework for robustness analysis of visual qa models,” in AAAI , 2019
2019
Later among the works it cites.
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai, “Vl-bert: Pre-training of generic visual-linguistic representations,” ICLR , 2020
2020
Closest in time.
K. Yi, C. Gan, Y. Li, P. Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum, “Clevrer: Collision events for video representation and reasoning,” ICLR , 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
B. Patro, S. Patel, and V. Namboodiri, “Robust explanations for visual question answering,” in WACV , 2020
2020
Closest in time.
V. Agarwal, R. Shetty, and M. Fritz, “Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing,” in CVPR , 2020
2020
Closest in time.
L. Chen, X. Yan, J. Xiao, H. Zhang, S. Pu, and Y. Zhuang, “Counterfactual samples synthesizing for robust visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 800–10 809
2020
Closest in time.