Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel Question-Guided Hybrid Convolution (QGHC) network for Visual Question Answering (VQA).
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., Hinton, G.E.: · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Frome, A., Corrado, G.S., Shlens, J., Bengio, S., Dean, J., Mikolov, T., et al.: · 2013
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., Le, Q.V.: · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., Bengio, Y.: · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., Parikh, D.: · 2015
Earlier work this paper cites.
Simple baseline for visual question answering
Zhou, B., Tian, Y., Sukhbaatar, S., Szlam, A., Fergus, R.: · 2015
Earlier work this paper cites.
Bilinear cnn models for fine-grained visual recognition
Lin, T.Y., RoyChowdhury, A., Maji, S.: · 2015
Earlier work this paper cites.
Abc-cnn: An attention based convolutional neural network for visual question answering
Chen, K., Wang, J., Chen, L.C., Gao, H., Xu, W., Nevatia, R.: · 2015
Earlier work this paper cites.
Skip-thought vectors
Kiros, R., Zhu, Y., Salakhutdinov, R.R., Zemel, R., Urtasun, R., Torralba, A., Fidler, S.: · 2015
Earlier work this paper cites.
Natural language object retrieval
Hu, R., Xu, H., Rohrbach, M., Feng, J., Saenko, K., Darrell, T.: · 2016
Earlier work this paper cites.
Learning deep representations of fine-grained visual descriptions
Reed, S., Akata, Z., Lee, H., Schiele, B.: · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., Rohrbach, M.: · 2016
Earlier work this paper cites.
Hadamard product for low-rank bilinear pooling
Kim, J.H., On, K.W., Kim, J., Ha, J.W., Zhang, B.T.: · 2016
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
Chollet, F.: · 2016
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., He, K.: · 2016
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Xu, H., Saenko, K.: · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Yang, Z., He, X., Gao, J., Deng, L., Smola, A.: · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
Lu, J., Yang, J., Batra, D., Parikh, D.: · 2016
Cited alongside, same era.
Image question answering using convolutional neural network with dynamic parameter prediction
Noh, H., Hongsuck Seo, P., Han, B.: · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Scene graph generation from objects, phrases and region captions
Li, Y., Ouyang, W., Zhou, B., Wang, K., Wang, X.: · 2017
Later among the works it cites.
Mutan: Multimodal tucker fusion for visual question answering
Ben-younes, H., Cadene, R., Cord, M., Thome, N.: · 2017
Later among the works it cites.
Modulating early visual processing by language
de Vries, H., Strub, F., Mary, J., Larochelle, H., Pietquin, O., Courville, A.: · 2017
Later among the works it cites.
Tracking by natural language specification
Li, Z., Tao, R., Gavves, E., Snoek, C.G., Smeulders, A., et al.: · 2017
Later among the works it cites.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
Zhang, X., Zhou, X., Lin, M., Sun, J.: · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
He, K., Zhang, X., Ren, S., Sun, J.: · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T., Kingma, D.P.: · 2016
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C.L., Girshick, R.: · 2016
Cited alongside, same era.
Learning deep features for discriminative localization
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., Torralba, A.: · 2016
Cited alongside, same era.
Neural module networks
Andreas, J., Rohrbach, M., Darrell, T., Klein, D.: · 2016
Cited alongside, same era.
Multimodal residual learning for visual qa
Kim, J.H., Lee, S.W., Kwak, D., Heo, M.O., Kim, J., Ha, J.W., Zhang, B.T.: · 2016
Cited alongside, same era.
Learning to compose neural networks for question answering
Andreas, J., Rohrbach, M., Darrell, T., Klein, D.: · 2016
Cited alongside, same era.
Identity-aware textual-visual matching with latent co-attention
Li Shuang, Xiao Tong, Li Hongsheng, Yang Wei, and Wang Xiaogang: · 2017
Later among the works it cites.
Person search with natural language description
Li Shuang, Xiao Tong, Li Hongsheng, Zhou Bolei, Yue Dayu, and Wang Xiaogang: · 2017
Later among the works it cites.
Learning to reason: End-to-end module networks for visual question answering
Hu, R., Andreas, J., Rohrbach, M., Darrell, T., Saenko, K.: · 2017
Later among the works it cites.
Inferring and executing programs for visual reasoning
Johnson, J., Hariharan, B., van der Maaten, L., Hoffman, J., Fei-Fei, L., Zitnick, C.L., Girshick, R.: · 2017
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
Perez, Ethan and Strub, Florian and De Vries, Harm and Dumoulin, Vincent and Courville, Aaron: · 2017
Later among the works it cites.
Co-attending Free-form Regions and Detections with Multi-modal Multiplicative Feature Embedding for Visual Question Answering
Lu, P., Li, H., Zhang, W., Wang, J., Wang, X.: · 2018
Closest in time.
R-VQA: Learning Visual Relation Facts with Semantic Attention for Visual Question Answering
Lu, P., Ji, L., Zhang, W., Duan, N., Zhou, M., Wang, J.: · 2018
Closest in time.
Diversity Regularized Spatiotemporal Attention for Video-based Person Re-identification
Li Shuang, Bak Slawomir, Carr Peter, and Wang Xiaogang: · 2018
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., Bengio, Y.: · 2057
Closest in time.