Fetching the paper…
Reading the bibliography…
Artificial Intelligence (AI) and its applications have sparked extraordinary interest in recent years.
1909
Earlier work this paper cites.
Z. S. Harris, “Distributional structure,” Word , vol. 10, no. 2-3, pp. 146–162, 1954
1954
Earlier work this paper cites.
A. W. Burks, D. W. Warren, and J. B. Wright, “An analysis of a logical machine using parenthesis-free notation,” Mathematical tables and other aids to computation , vol. 8, no. 46, pp. 53–57, 1954
1954
Earlier work this paper cites.
H. R. Burke, “Raven’s progressive matrices: A review and critical evaluation,” The Journal of Genetic Psychology , vol. 93, no. 2, pp. 199–228, 1958
1958
Earlier work this paper cites.
L. R. Tucker, “Some mathematical notes on three-mode factor analysis,” Psychometrika , vol. 31, no. 3, pp. 279–311, 1966
1966
Earlier work this paper cites.
T. Kasami, “An efficient recognition and syntax-analysis algorithm for context-free languages,” Coordinated Science Laboratory Report no. R-257 , 1966
1966
Earlier work this paper cites.
Z. Wu and M. Palmer, “Verb semantics and lexical selection,” arXiv preprint cmp-lg/9406033 , 1994
1994
Earlier work this paper cites.
Y. LeCun, Y. Bengio et al. , “Convolutional networks for images, speech, and time series,” The handbook of brain theory and neural networks , vol. 3361, no. 10, p. 1995, 1995
1995
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
R. Cadene, H. Ben-Younes, M. Cord, and N. Thome, “Murel: Multimodal relational reasoning for visual question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 1989–1998
1998
Earlier work this paper cites.
H. Liu and P. Singh, “Conceptnet—a practical commonsense reasoning tool-kit,” BT technology journal , vol. 22, no. 4, pp. 211–226, 2004
2004
Earlier work this paper cites.
M. Rickert, M. E. Foster, M. Giuliani, T. By, G. Panin, and A. Knoll, “Integrating language, vision and action for human robot dialog systems,” in International Conference on Universal Access in Human-Computer Interaction . Springer, 2007, pp. 987–995
2007
Earlier work this paper cites.
S. Auer, C. Bizer, G. Kobilarov, J. Lehmann, R. Cyganiak, and Z. Ives, “Dbpedia: A nucleus for a web of open data,” in The semantic web . Springer, 2007, pp. 722–735
2007
Earlier work this paper cites.
A. Baumann, M. Boltz, J. Ebling, M. Koenig, H. Loos, M. Merkel, W. Niem, J. Warzelhan, and J. Yu, “A review and comparison of measures for automatic video surveillance systems,” EURASIP Journal on Image and Video Processing , vol. 2008, pp. 1–30, 2008
2008
Earlier work this paper cites.
K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor, “Freebase: a collaboratively created graph database for structuring human knowledge,” in Proceedings of the 2008 ACM SIGMOD international conference on Management of data , 2008, pp. 1247–1250
2008
Earlier work this paper cites.
O. Etzioni, M. Banko, S. Soderland, and D. S. Weld, “Open information extraction from the web,” Communications of the ACM , vol. 51, no. 12, pp. 68–74, 2008
2008
Earlier work this paper cites.
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, P.-A. Manzagol, and L. Bottou, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion.” Journal of machine learning research , vol. 11, no. 12, 2010
2010
Earlier work this paper cites.
A. Carlson, J. Betteridge, B. Kisiel, B. Settles, E. R. Hruschka, and T. M. Mitchell, “Toward an architecture for never-ending language learning,” in Twenty-Fourth AAAI conference on artificial intelligence , 2010
2010
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal processing magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “Indoor segmentation and support inference from rgbd images,” in European conference on computer vision . Springer, 2012, pp. 746–760
2012
Earlier work this paper cites.
J. Hoffart, F. M. Suchanek, K. Berberich, and G. Weikum, “Yago2: A spatially and temporally enhanced knowledge base from wikipedia,” Artificial Intelligence , vol. 194, pp. 28–61, 2013
2013
Earlier work this paper cites.
M. Malinowski and M. Fritz, “A multi-world approach to question answering about real-world scenes based on uncertain input,” Advances in neural information processing systems , vol. 27, pp. 1682–1690, 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European conference on computer vision . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
N. Tandon, G. De Melo, and G. Weikum, “Acquiring comparative commonsense knowledge from the web,” in Twenty-Eighth AAAI Conference on Artificial Intelligence , 2014
2014
Earlier work this paper cites.
N. Tandon, G. De Melo, F. Suchanek, and G. Weikum, “Webchild: Harvesting and organizing commonsense knowledge from the web,” in Proceedings of the 7th ACM international conference on Web search and data mining , 2014, pp. 523–532
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
N. Liu, G. Gimel’farb, and P. Delmas, “High-order mgrf models for contrast/offset invariant texture retrieval,” in Proceedings of the 29th International Conference on Image and Vision Computing New Zealand , 2014, pp. 96–101
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 1532–1543
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Andreas and D. Klein, “How much do word embeddings encode about syntax?” in Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , 2014, pp. 822–827
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 2625–2634
2015
Earlier work this paper cites.
M. Ren, R. Kiros, and R. Zemel, “Exploring models and data for image question answering,” Advances in neural information processing systems , vol. 28, pp. 2953–2961, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 2425–2433
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 1–9
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3156–3164
2015
Earlier work this paper cites.
M. Ren, R. Kiros, and R. Zemel, “Image question answering: A visual semantic embedding model and a new dataset,” Proc. Advances in Neural Inf. Process. Syst , vol. 1, no. 2, p. 5, 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: towards real-time object detection with region proposal networks,” IEEE transactions on pattern analysis and machine intelligence , vol. 39, no. 6, pp. 1137–1149, 2016
2016
Earlier work this paper cites.
R. Bernardi, R. Cakici, D. Elliott, A. Erdem, E. Erdem, N. Ikizler-Cinbis, F. Keller, A. Muscat, and B. Plank, “Automatic description generation from images: A survey of models, datasets, and evaluation measures,” Journal of Artificial Intelligence Research , vol. 55, pp. 409–442, 2016
2016
Earlier work this paper cites.
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei, “Visual7W: Grounded Question Answering in Images,” in IEEE Conference on Computer Vision and Pattern Recognition , 2016
2016
Earlier work this paper cites.
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh, “Yin and yang: Balancing and answering binary visual questions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 5014–5022
2016
Earlier work this paper cites.
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei, “Visual7w: Grounded question answering in images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4995–5004
2016
Earlier work this paper cites.
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein, “Neural module networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 39–48
2016
Earlier work this paper cites.
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi, “A diagram is worth a dozen images,” in European conference on computer vision . Springer, 2016, pp. 235–251
2016
Earlier work this paper cites.
H. Noh, P. H. Seo, and B. Han, “Image question answering using convolutional neural network with dynamic parameter prediction,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 30–38
2016
Earlier work this paper cites.
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola, “Stacked attention networks for image question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 21–29
2016
Cited alongside, same era.
H. Xu and K. Saenko, “Ask, attend and answer: Exploring question-guided spatial attention for visual question answering,” in European Conference on Computer Vision . Springer, 2016, pp. 451–466
2016
Cited alongside, same era.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” Advances in neural information processing systems , vol. 29, pp. 289–297, 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
W. Norcliffe-Brown, S. Vafeias, and S. Parisot, “Learning conditioned graph structures for interpretable visual question answering,” Advances in neural information processing systems , vol. 31, 2018
2018
Later among the works it cites.
D.-K. Nguyen and T. Okatani, “Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 6087–6096
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
C. Xiong, S. Merity, and R. Socher, “Dynamic memory networks for visual and textual question answering,” in International conference on machine learning . PMLR, 2016, pp. 2397–2406
2016
Cited alongside, same era.
Q. Wu, P. Wang, C. Shen, A. Dick, and A. Van Den Hengel, “Ask me anything: Free-form visual question answering based on knowledge from external sources,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4622–4630
2016
Cited alongside, same era.
A. Jabri, A. Joulin, and L. Van Der Maaten, “Revisiting visual question answering baselines,” in European conference on computer vision . Springer, 2016, pp. 727–739
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
Y. Gao, O. Beijbom, N. Zhang, and T. Darrell, “Compact bilinear pooling,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 317–326
2016
Cited alongside, same era.
2016
Cited alongside, same era.
A. Kumar, O. Irsoy, P. Ondruska, M. Iyyer, J. Bradbury, I. Gulrajani, V. Zhong, R. Paulus, and R. Socher, “Ask me anything: Dynamic memory networks for natural language processing,” in International conference on machine learning . PMLR, 2016, pp. 1378–1387
2016
Cited alongside, same era.
M. Z. Hossain, F. Sohel, M. F. Shiratuddin, and H. Laga, “A comprehensive survey of deep learning for image captioning,” ACM Computing Surveys (CsUR) , vol. 51, no. 6, pp. 1–36, 2019
2019
Later among the works it cites.
D. Liu, M. Bober, and J. Kittler, “Visual semantic information pursuit: A survey,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 4, pp. 1404–1422, 2019
2019
Later among the works it cites.
D. A. Hudson and C. D. Manning, “Gqa: A new dataset for real-world visual reasoning and compositional question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6700–6709
2019
Later among the works it cites.
R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi, “From recognition to cognition: Visual commonsense reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6720–6731
2019
Later among the works it cites.
C. Zhang, F. Gao, B. Jia, Y. Zhu, and S.-C. Zhu, “Raven: A dataset for relational and analogical visual reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 5317–5327
2019
Later among the works it cites.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “Ok-vqa: A visual question answering benchmark requiring external knowledge,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 3195–3204
2019
Later among the works it cites.
S. Shah, A. Mishra, N. Yadati, and P. P. Talukdar, “Kvqa: Knowledge-aware visual question answering,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 8876–8884
2019
Later among the works it cites.
A. Suhr and Y. Artzi, “Nlvr2 visual bias analysis,” arXiv preprint arXiv:1909.10411 , 2019
2019
Later among the works it cites.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8317–8326
2019
Later among the works it cites.
W. A. Qader, M. M. Ameen, and B. I. Ahmed, “An overview of bag of words; importance, implementation, applications, and challenges,” in 2019 International Engineering Conference (IEC) . IEEE, 2019, pp. 200–204
2019
Later among the works it cites.
2019
Later among the works it cites.
H. Ben-Younes, R. Cadene, N. Thome, and M. Cord, “Block: Bilinear superdiagonal fusion for visual question answering and visual relationship detection,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 8102–8109
2019
Later among the works it cites.
Z. Yu, J. Yu, Y. Cui, D. Tao, and Q. Tian, “Deep modular co-attention networks for visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6281–6290
2019
Later among the works it cites.
L. Li, Z. Gan, Y. Cheng, and J. Liu, “Relation-aware graph attention network for visual question answering,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 10 313–10 322
2019
Later among the works it cites.
2019
Later among the works it cites.
K. Chen, J. Pang, J. Wang, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Shi, W. Ouyang et al. , “Hybrid task cascade for instance segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 4974–4983
2019
Later among the works it cites.
M. Sridharan and T. Swapna, “Amrita school of engineering-cse at semeval-2019 task 6: Manipulating attention with temporal convolutional neural network for offense identification and classification,” in Proceedings of the 13th International Workshop on Semantic Evaluation , 2019, pp. 540–546
2019
Later among the works it cites.
J. Shi, H. Zhang, and J. Li, “Explainable and explicit visual reasoning over scene graphs,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8376–8384
2019
Later among the works it cites.
2019
Later among the works it cites.
P. Xiong, H. Zhan, X. Wang, B. Sinha, and Y. Wu, “Visual query answering by entity-attribute graph matching and reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8357–8366
2019
Later among the works it cites.
P. Gao, Z. Jiang, H. You, P. Lu, S. C. Hoi, X. Wang, and H. Li, “Dynamic fusion with intra-and inter-modality attention flow for visual question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6639–6648
2019
Later among the works it cites.
2020
Later among the works it cites.
J. Hong, S. Park, and H. Byun, “Selective residual learning for visual question answering,” Neurocomputing , vol. 402, pp. 366–374, 2020
2020
Later among the works it cites.
M. H. Vu, T. Löfstedt, T. Nyholm, and R. Sznitman, “A question-centric model for visual question answering in medical imaging,” IEEE transactions on medical imaging , vol. 39, no. 9, pp. 2856–2868, 2020
2020
Later among the works it cites.
Y. Liu, X. Zhang, Z. Zhao, B. Zhang, L. Cheng, and Z. Li, “Alsa: Adversarial learning of supervised attentions for visual question answering,” IEEE Transactions on Cybernetics , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
H. Zhong, J. Chen, C. Shen, H. Zhang, J. Huang, and X.-S. Hua, “Self-adaptive neural module transformer for visual question answering,” IEEE Transactions on Multimedia , vol. 23, pp. 1264–1273, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Yu, Z. Zhu, Y. Wang, W. Zhang, Y. Hu, and J. Tan, “Cross-modal knowledge reasoning for knowledge-based visual question answering,” Pattern Recognition , vol. 108, p. 107563, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
X. Zhu, Z. Mao, Z. Chen, Y. Li, Z. Wang, and B. Wang, “Object-difference drived graph convolutional networks for visual question answering,” Multimedia Tools and Applications , pp. 1–19, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Khademi, “Multimodal neural graph memory networks for visual question answering,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 7177–7188
2020
Later among the works it cites.
2020
Later among the works it cites.
R. Y. Zakari, Z. K. Lawal, and I. Abdulmumin, “A systematic literature review of hausa natural language processing,” International Journal of Computer and Information Technology (2279-0764) , vol. 10, no. 4, 2021
2021
Later among the works it cites.
S. H. Hosseinabad, M. Safayani, and A. Mirzaei, “Multiple answers to a question: a new approach for visual question answering,” The Visual Computer , vol. 37, no. 1, pp. 119–131, 2021
2021
Later among the works it cites.
Q. Cao, B. Li, X. Liang, K. Wang, and L. Lin, “Knowledge-routed visual question reasoning: Challenges for deep representation embedding,” IEEE Transactions on Neural Networks and Learning Systems , 2021
2021
Later among the works it cites.
W. Zhang, J. Yu, Y. Wang, and W. Wang, “Multimodal deep fusion for image question answering,” Knowledge-Based Systems , vol. 212, p. 106639, 2021
2021
Later among the works it cites.
Z. Ma, W. Zheng, X. Chen, and L. Yin, “Joint embedding vqa model based on dynamic word vector,” PeerJ Computer Science , vol. 7, p. e353, 2021
2021
Later among the works it cites.
D. Song, S. Ma, Z. Sun, S. Yang, and L. Liao, “Kvl-bert: Knowledge enhanced visual-and-linguistic bert for visual commonsense reasoning,” Knowledge-Based Systems , vol. 230, p. 107408, 2021
2021
Later among the works it cites.
K. Marino, X. Chen, D. Parikh, A. Gupta, and M. Rohrbach, “Krisp: Integrating implicit and symbolic knowledge for open-domain knowledge-based vqa,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 111–14 121
2021
Later among the works it cites.
2021
Later among the works it cites.
R. Dua, S. S. Kancheti, and V. N. Balasubramanian, “Beyond vqa: Generating multi-word answers and rationales to visual questions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 1623–1632
2021
Later among the works it cites.
D. Guo, C. Xu, and D. Tao, “Bilinear graph networks for visual question answering,” IEEE Transactions on Neural Networks and Learning Systems , 2021
2021
Later among the works it cites.
W. Zhang, J. Yu, W. Zhao, and C. Ran, “Dmrfnet: Deep multimodal reasoning and fusion for visual question answering and explanation generation,” Information Fusion , vol. 72, pp. 70–79, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.