Fetching the paper…
Reading the bibliography…
Learning an effective attention mechanism for multimodal data is important in many vision-and-language tasks that require a synergic understanding of both the visual and textual contents.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
Z. Yu, F. Wu, Y. Yang, Q. Tian, J. Luo, and Y. Zhuang, “Discriminative coupled dictionary hashing for fast cross-media retrieval,” in International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) , 2014, pp. 395–404
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Mnih, N. Heess, A. Graves et al. , “Recurrent models of visual attention,” in NIPS , 2014, pp. 2204–2212
2014
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2014, pp. 787–798
2014
Earlier work this paper cites.
C. L. Zitnick and P. Dollár, “Edge boxes: Locating object proposals from edges,” in European Conference on Computer Vision (ECCV) , 2014, pp. 391–405
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation.” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , vol. 14, 2014, pp. 1532–1543
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European Conference on Computer Vision (ECCV) , 2014, pp. 740–755
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention.” in International Conference on Machine Learning (ICML) , vol. 14, 2015, pp. 77–81
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh, “Vqa: Visual question answering,” in IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 2425–2433
2015
Earlier work this paper cites.
K. Gregor, I. Danihelka, A. Graves, D. J. Rezende, and D. Wierstra, “Draw: A recurrent neural network for image generation,” in International Conference on Machine Learning (ICML) , 2015, pp. 1462–1471
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
R. Girshick, “Fast r-cnn,” in IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 1440–1448
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in NIPS , 2015, pp. 91–99
2015
Earlier work this paper cites.
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele, “Grounding of textual phrases in images by reconstruction,” in European Conference on Computer Vision (ECCV) , 2016, pp. 817–834
2016
Earlier work this paper cites.
L.-C. Chen, Y. Yang, J. Wang, W. Xu, and A. L. Yuille, “Attention to scale: Scale-aware semantic image segmentation,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 3640–3649
2016
Earlier work this paper cites.
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola, “Stacked attention networks for image question answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 21–29
2016
Earlier work this paper cites.
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach, “Multimodal compact bilinear pooling for visual question answering and visual grounding,” Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2016
2016
Earlier work this paper cites.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” in NIPS , 2016, pp. 289–297
2016
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in IEEE International Conference on Computer Vision (ICCV) , 2016, pp. 11–20
2016
Cited alongside, same era.
2016
Cited alongside, same era.
K. J. Shih, S. Singh, and D. Hoiem, “Where to look: Focus regions for visual question answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 4613–4621
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Cited alongside, same era.
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko, “Modeling relationships in referential expressions with compositional modular networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 1115–1124
2017
Later among the works it cites.
E. Yang, C. Deng, C. Li, W. Liu, J. Li, and D. Tao, “Shared predictive cross-modal deep quantization,” IEEE transactions on neural networks and learning systems , vol. 29, no. 11, pp. 5292–5303, 2018
2018
Later among the works it cites.
J. Song, Y. Guo, L. Gao, X. Li, A. Hanjalic, and H. T. Shen, “From deterministic to generative: multi-modal stochastic rnns for video captioning,” IEEE transactions on neural networks and learning systems , 2018
2018
Later among the works it cites.
J.-H. Kim, J. Jun, and B.-T. Zhang, “Bilinear attention networks,” NIPS , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2016
Cited alongside, same era.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in European Conference on Computer Vision (ECCV) . Springer, 2016, pp. 69–85
2016
Cited alongside, same era.
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in European Conference on Computer Vision (ECCV) . Springer, 2016, pp. 21–37
2016
Cited alongside, same era.
T. Dozat and C. D. Manning, “Deep biaffine attention for neural dependency parsing,” in International Conference on Learning Representations (ICLR) , 2017. [Online]. Available: https://nlp.stanford.edu/pubs/dozat2017deep.pdf
2017
Cited alongside, same era.
Z. Yu, J. Yu, J. Fan, and D. Tao, “Multi-modal factorized bilinear pooling with co-attention learning for visual question answering,” IEEE International Conference on Computer Vision (ICCV) , pp. 1839–1848, 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017, pp. 6000–6010
2017
Cited alongside, same era.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Cited alongside, same era.
D.-K. Nguyen and T. Okatani, “Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 7794–7803
2018
Later among the works it cites.
H. Hu, J. Gu, Z. Zhang, J. Dai, and Y. Wei, “Relation networks for object detection,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Later among the works it cites.
Z. Yu, J. Yu, C. Xiang, J. Fan, and D. Tao, “Beyond bilinear: Generalized multi-modal factorized high-order pooling for visual question answering,” IEEE Transactions on Neural Networks and Learning Systems , 2018
2018
Later among the works it cites.
Z. Yu, J. Yu, C. Xiang, Z. Zhao, Q. Tian, and D. Tao, “Rethinking diversified and discriminative proposal generation for visual grounding,” International Joint Conference on Artificial Intelligence (IJCAI) , 2018
2018
Later among the works it cites.
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg, “Mattnet: Modular attention network for referring expression comprehension,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 1307–1315
2018
Later among the works it cites.
B. Zhuang, Q. Wu, C. Shen, I. Reid, and A. van den Hengel, “Parallel attention: A unified framework for visual object discovery through dialogs and queries,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 4252–4261
2018
Later among the works it cites.
C. Deng, Q. Wu, Q. Wu, F. Hu, F. Lyu, and M. Tan, “Visual grounding via accumulated attention,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 7746–7755
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” URL https://s3-us-west-2. amazonaws. com/openai-assets/research-covers/languageunsupervised/language understanding paper. pdf , 2018
2018
Later among the works it cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
Y. Zhang, J. Hare, and A. Prügel-Bennett, “Learning to count objects in natural images for visual question answering,” International Conference on Learning Representation (ICLR) , 2018
2018
Later among the works it cites.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
H. Zhang, Y. Niu, and S.-F. Chang, “Grounding referring expressions in images by variational context,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
Later among the works it cites.
C. Zheng, L. Pan, and P. Wu, “Multimodal deep network embedding with integrated structure and attribute information,” IEEE transactions on neural networks and learning systems , 2019
2019
Closest in time.
X. Li, J. Song, L. Gao, X. Liu, W. Huang, X. He, and C. Gan, “Beyond rnns: Positional self-attention with co-attention for video question answering,” in AAAI , 2019
2019
Closest in time.
Z. Yu, J. Yu, Y. Cui, D. Tao, and Q. Tian, “Deep modular co-attention networks for visual question answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 6281–6290
2019
Closest in time.