Fetching the paper…
Reading the bibliography…
Answering questions according to multi-modal context is a challenging problem as it requires a deep integration of different data sources.
H. Liu and P. Singh, “Conceptnet?a practical commonsense reasoning tool-kit,” BT technology journal , vol. 22, no. 4, pp. 211–226, 2004
2004
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh, “Vqa: Visual question answering,” in ICCV , 2015
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual Question Answering,” in ICCV , 2015
2015
Earlier work this paper cites.
L. Ma, Z. Lu, L. Shang, and H. Li, “Multimodal convolutional neural networks for matching image and sentence,” in ICCV , 2015
2015
Earlier work this paper cites.
M. Malinowski, M. Rohrbach, and M. Fritz, “Ask your neurons: A neural-based approach to answering questions about images,” in ICCV , 2015
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” 2015
2015
Earlier work this paper cites.
S. Sukhbaatar, J. Weston, R. Fergus et al. , “End-to-end memory networks,” in NIPS , 2015
2015
Earlier work this paper cites.
Y. Gao, O. Beijbom, N. Zhang, and T. Darrell, “Compact bilinear pooling,” in CVPR , 2016
2016
Earlier work this paper cites.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” in NIPS , 2016
2016
Earlier work this paper cites.
P. Pan, Z. Xu, Y. Yang, F. Wu, and Y. Zhuang, “Hierarchical recurrent neural encoder for video representation with application to captioning,” in CVPR , 2016
2016
Earlier work this paper cites.
K. J. Shih, S. Singh, and D. Hoiem, “Where to look: Focus regions for visual question answering,” in CVPR , 2016
2016
Earlier work this paper cites.
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler, “MovieQA: Understanding Stories in Movies through Question-Answering,” in CVPR , 2016
2016
Earlier work this paper cites.
C. Xiong, S. Merity, and R. Socher, “Dynamic memory networks for visual and textual question answering,” in ICML , 2016
2016
Cited alongside, same era.
H. Xu and K. Saenko, “Ask, attend and answer: Exploring question-guided spatial attention for visual question answering,” in ECCV , 2016
2016
Cited alongside, same era.
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo, “Image captioning with semantic attention,” in CVPR , 2016
2016
Cited alongside, same era.
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei, “Visual7w: Grounded question answering in images,” in CVPR , 2016
2016
Cited alongside, same era.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in CVPR , 2017
2017
Cited alongside, same era.
S. Na, S. Lee, J. Kim, and G. Kim, “A read-write memory network for movie story understanding,” in CVPR , 2017
2017
Later among the works it cites.
F. Nian, T. Li, Y. Wang, X. Wu, B. Ni, and C. Xu, “Learning explicit video attributes from mid-level representation for video captioning,” Computer Vision and Image Understanding , vol. 163, pp. 126–138, 2017
2017
Later among the works it cites.
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in CVPR , 2017
2017
Later among the works it cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,” in CVPR , 2017
2017
Cited alongside, same era.
J. Hu, D. Fan, S. Yao, and J. Oh, “Answer-aware attention on grounded question answering in images,” 2017
2017
Cited alongside, same era.
Y. Huang, W. Wang, and L. Wang, “Instance-aware image and sentence matching with selective multimodal lstm,” in CVPR , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick, “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,” in CVPR , 2017
2017
Cited alongside, same era.
A. Kembhavi, M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi, “Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,” in CVPR , 2017
2017
Cited alongside, same era.
K.-M. Kim, M.-O. Heo, S.-H. Choi, and B.-T. Zhang, “Deepstory: video story qa by deep embedded memory networks,” in IJCAI , 2017
2017
Cited alongside, same era.
2018
Closest in time.
U. Jain, S. Lazebnik, and A. Schwing, “Two can play this game: Visual dialog with discriminative question generation and answering,” in CVPR , 2018
2018
Closest in time.
J. Liang, L. Jiang, L. Cao, L.-J. Li, and A. Hauptmann, “Focal visual-text attention for visual question answering,” in CVPR , 2018
2018
Closest in time.
B. Wang, Y. Xu, Y. Han, and R. Hong, “Movie question answering: Remembering the textual cues for layered visual contents,” in AAAI , 2018
2018
Closest in time.
L. Wang, Y. Li, J. Huang, and S. Lazebnik, “Learning two-branch neural networks for image-text matching tasks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2018
2018
Closest in time.
Q. Wu, C. Shen, P. Wang, A. Dick, and A. van den Hengel, “Image captioning and visual question answering based on attributes and external knowledge,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 6, pp. 1367–1381, 2018
2018
Closest in time.
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, “End-to-end dense video captioning with masked transformer,” in CVPR , 2018
2018
Closest in time.
M. Zhang, Y. Yang, H. Zhang, Y. Ji, H. T. Shen, and T.-S. Chua, “More is better: Precise and detailed image captioning using online positive recall and missing concepts mining,” IEEE Transactions on Image Processing , vol. 28, no. 1, pp. 32–44, 2019
2019
Closest in time.