Fetching the paper…
Reading the bibliography…
The Visual Question Answering (VQA) task combines challenges for processing data with both Visual and Linguistic processing, to answer basic `common sense' questions about given images.
Neural computation 9
Hochreiter, S., Schmidhuber, J.: Long short-term memory · 1997
Earlier work this paper cites.
Design and Applications 5
Medsker, L.R., Jain, L.: Recurrent neural networks · 2001
Earlier work this paper cites.
In: NIPS, pp. 1097–1105 (2012)
Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks · 2012
Earlier work this paper cites.
In: ECCV, pp. 746–760 (2012)
Silberman, N., Hoiem, D., Kohli, P., Fergus, R.: Indoor segmentation and support inference from rgbd images · 2012
Earlier work this paper cites.
In: IEEE CVPR, pp. 564–571 (2013)
Gupta, S., Arbelaez, P., Malik, J.: Perceptual organization and recognition of indoor scenes from rgb-d images · 2013
Earlier work this paper cites.
arXiv preprint arXiv:1412.3555 (2014)
Chung, J., Gulcehre, C., Cho, K., Bengio, Y.: Empirical evaluation of gated recurrent neural networks on sequence modeling · 2014
Earlier work this paper cites.
In: ECCV, pp. 740–755 (2014)
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context · 2014
Earlier work this paper cites.
In: NIPS, pp. 1682–1690 (2014)
Malinowski, M., Fritz, M.: A multi-world approach to question answering about real-world scenes based on uncertain input · 2014
Earlier work this paper cites.
In: EMNLP, pp. 1532–1543 (2014)
Pennington, J., Socher, R., Manning, C.: Glove: Global vectors for word representation · 2014
Earlier work this paper cites.
arXiv preprint arXiv:1409.1556 (2014)
Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition · 2014
Earlier work this paper cites.
In: IEEE ICCV, pp. 2425–2433 (2015)
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., Parikh, D.: Vqa: Visual question answering · 2015
Earlier work this paper cites.
In: Advances in neural information processing systems, pp. 2953–2961 (2015)
Ren, M., Kiros, R., Zemel, R.: Exploring models and data for image question answering · 2015
Earlier work this paper cites.
In: NIPS, pp. 91–99 (2015)
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks · 2015
Earlier work this paper cites.
In: IEEE ICCV, pp. 2461–2469 (2015)
Yu, L., Park, E., Berg, A.C., Berg, T.L.: Visual madlibs: Fill in the blank description generation and question answering · 2015
Cited alongside, same era.
In: IEEE CVPR, pp. 770–778 (2016)
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition · 2016
Cited alongside, same era.
In: IEEE CVPR, pp. 2818–2826 (2016)
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., Wojna, Z.: Rethinking the inception architecture for computer vision · 2016
Cited alongside, same era.
In: IEEE CVPR, pp. 4631–4640 (2016)
Tapaswi, M., Zhu, Y., Stiefelhagen, R., Torralba, A., Urtasun, R., Fidler, S.: Movieqa: Understanding stories in movies through question-answering · 2016
Cited alongside, same era.
In: IEEE CVPR, pp. 21–29 (2016)
Yang, Z., He, X., Gao, J., Deng, L., Smola, A.: Stacked attention networks for image question answering · 2016
Cited alongside, same era.
In: IEEE CVPR, pp. 7680–7688 (2018)
Patro, B., Namboodiri, V.P.: Differential attention for visual question answering · 2018
Later among the works it cites.
In: IEEE CVPR, pp. 4223–4232 (2018)
Teney, D., Anderson, P., He, X., van den Hengel, A.: Tips and tricks for visual question answering: Learnings from the 2017 challenge · 2018
Later among the works it cites.
In: NIPS, pp. 1031–1042 (2018)
Yi, K., Wu, J., Gan, C., Torralba, A., Kohli, P., Tenenbaum, J.: Neural-symbolic vqa: Disentangling reasoning from vision and language understanding · 2018
Later among the works it cites.
In: AAAI (2019)
Shah, S., Mishra, A., Yadati, N., Talukdar, P.P.: Kvqa: Knowledge-aware visual question answering · 2019
Closest in time.
AAAI 2019 (2019)
Wu, C., Liu, J., Wang, X., Li, R.: Differential networks for visual question answering · 2019
Closest in time.
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6669–6678 (2019)
Zheng, Z., Wang, W., Qi, S., Zhu, S.C.: Reasoning visual dialogs with structural and partial observations · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhu, Y., Groth, O., Bernstein, M., Fei-Fei, L.: Visual7w: Grounded question answering in images · 2016
Cited alongside, same era.
arXiv preprint arXiv:1708.01336 (2017)
Jiang, L., Liang, J., Cao, L., Kalantidis, Y., Farfade, S., Hauptmann, A.: Memexqa: Visual memex question answering · 2017
Cited alongside, same era.
In: IEEE CVPR, pp. 2901–2910 (2017)
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., Girshick, R.: Clevr: A diagnostic dataset for compositional language and elementary visual reasoning · 2017
Cited alongside, same era.
In: ICCV (2017)
Kafle, K., Kanan, C.: An analysis of visual question answering algorithms · 2017
Cited alongside, same era.
arXiv preprint arXiv:1810.12440 (2018)
Acharya, M., Kafle, K., Kanan, C.: Tallyqa: Answering complex counting questions · 2018
Cited alongside, same era.
arXiv preprint arXiv:1807.09956 (2018)
Jiang, Y., Natarajan, V., Chen, X., Rohrbach, M., Batra, D., Parikh, D.: Pythia v0. 1: the winning entry to the vqa challenge 2018 · 2018
Cited alongside, same era.
arXiv preprint arXiv:1807.09956 (2018)
Jiang, Y., Natarajan, V., Chen, X., Rohrbach, M., Batra, D., Parikh, D.: Pythia v0. 1: the winning entry to the vqa challenge 2018 · 2018
Cited alongside, same era.
Closest in time.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10,800–10,809 (2020)
Chen, L., Yan, X., Xiao, J., Zhang, H., Pu, S., Zhuang, Y.: Counterfactual samples synthesizing for robust visual question answering · 2020
Closest in time.
In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7166–7176 (2020)
Huang, Q., Wei, J., Cai, Y., Zheng, C., Chen, J., Leung, H.f., Li, Q.: Aligned dual channel graph convolutional network for visual question answering · 2020
Closest in time.
In: Proceedings of the 28th ACM International Conference on Multimedia, pp. 1227–1235 (2020)
Li, G., Wang, X., Zhu, W.: Boosting visual question answering with context-aware knowledge aggregation · 2020
Closest in time.
Pattern Recognition Letters 133
Li, W., Sun, J., Liu, G., Zhao, L., Fang, X.: Visual question answering with attention transfer and a cross-modal gating mechanism · 2020
Closest in time.
IEEE Transactions on Geoscience and Remote Sensing (2020)
Lobry, S., Marcos, D., Murray, J., Tuia, D.: Rsvqa: Visual question answering for remote sensing data · 2020
Closest in time.
Pattern Recognition 108
Yu, J., Zhu, Z., Wang, Y., Zhang, W., Hu, Y., Tan, J.: Cross-modal knowledge reasoning for knowledge-based visual question answering · 2020
Closest in time.
In: Proceedings of the 28th ACM International Conference on Multimedia, pp. 2345–2354 (2020)
Zhan, L.M., Liu, B., Fan, L., Chen, J., Wu, X.M.: Medical visual question answering via conditional reasoning · 2020
Closest in time.