Fetching the paper…
Reading the bibliography…
Visual question answering (VQA) is challenging because it requires a simultaneous understanding of both visual content of images and textual content of questions.
J. B. Tenenbaum and W. T. Freeman, “Separating style and content,” NIPS , pp. 662–668, 1997
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
S. Rendle, “Factorization machines,” in ICDM , 2010, pp. 995–1000
2010
Earlier work this paper cites.
F. Perronnin, J. Sánchez, and T. Mensink, “Improving the fisher kernel for large-scale image classification,” ECCV , pp. 143–156, 2010
2010
Earlier work this paper cites.
X. Geng, C. Yin, and Z.-H. Zhou, “Facial age estimation by learning from label distributions,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 35, no. 10, pp. 2401–2412, 2013
2013
Earlier work this paper cites.
F. Wu, Z. Yu, Y. Yang, S. Tang, Y. Zhang, and Y. Zhuang, “Sparse multi-modal hashing,” IEEE Transactions on Multimedia , vol. 16, no. 2, pp. 427–439, 2014
2014
Earlier work this paper cites.
Z. Yu, F. Wu, Y. Yang, Q. Tian, J. Luo, and Y. Zhuang, “Discriminative coupled dictionary hashing for fast cross-media retrieval,” in ACM SIGIR , 2014, pp. 395–404
2014
Earlier work this paper cites.
M. Malinowski and M. Fritz, “A multi-world approach to question answering about real-world scenes based on uncertain input,” in NIPS , 2014, pp. 1682–1690
2014
Earlier work this paper cites.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting.” JMLR , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV , 2014, pp. 740–755
2014
Earlier work this paper cites.
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” in ACM Multimedia , 2014, pp. 675–678
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation.” in EMNLP , vol. 14, 2014, pp. 1532–1543
2014
Earlier work this paper cites.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR , 2015, pp. 2625–2634
2015
Earlier work this paper cites.
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention.” in ICML , vol. 14, 2015, pp. 77–81
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh, “Vqa: Visual question answering,” in ICCV , 2015, pp. 2425–2433
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
T.-Y. Lin, A. RoyChowdhury, and S. Maji, “Bilinear cnn models for fine-grained visual recognition,” in ICCV , 2015, pp. 1449–1457
2015
Earlier work this paper cites.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
M. Malinowski, M. Rohrbach, and M. Fritz, “Ask your neurons: A neural-based approach to answering questions about images,” in ICCV , 2015, pp. 1–9
2015
Cited alongside, same era.
D. Tao, Y. Guo, M. Song, Y. Li, Z. Yu, and Y. Y. Tang, “Person re-identification by dual-regularized kiss metric learning,” IEEE Transactions on Image Processing , vol. 25, no. 6, pp. 2726–2738, 2016
2016
Later among the works it cites.
D. Tao, L. Jin, Y. Yuan, and Y. Xue, “Ensemble manifold rank preserving for acceleration-based human activity recognition,” IEEE Transactions on Neural Networks and Learning Systems , vol. 27, no. 6, pp. 1392–1404, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
H. Noh, P. Hongsuck Seo, and B. Han, “Image question answering using convolutional neural network with dynamic parameter prediction,” in CVPR , 2016, pp. 30–38
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Cited alongside, same era.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” in NIPS , 2016, pp. 289–297
2016
Cited alongside, same era.
J.-H. Kim, S.-W. Lee, D. Kwak, M.-O. Heo, J. Kim, J.-W. Ha, and B.-T. Zhang, “Multimodal residual learning for visual qa,” in NIPS , 2016, pp. 361–369
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
H. Xu and K. Saenko, “Ask, attend and answer: Exploring question-guided spatial attention for visual question answering,” in ECCV , 2016, pp. 451–466
2016
Later among the works it cites.
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein, “Neural module networks,” in CVPR , 2016, pp. 39–48
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
X. Shen, W. Liu, I. W. Tsang, Q.-S. Sun, and Y.-S. Ong, “Multilabel prediction via cross-view search,” IEEE Transactions on Neural Networks and Learning Systems , 2017
2017
Closest in time.
J.-H. Kim, K. W. On, W. Lim, J. Kim, J.-W. Ha, and B.-T. Zhang, “Hadamard Product for Low-rank Bilinear Pooling,” in ICLR , 2017
2017
Closest in time.
2017
Closest in time.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017, pp. 6000–6010
2017
Closest in time.
2017
Closest in time.
X. Shen, X. Tian, T. Liu, F. Xu, and D. Tao, “Continuous dropout,” IEEE Transactions on Neural Networks and Learning Systems , 2017
2017
Closest in time.
C. Shi, Z. Liu, X. Dong, and Y. Chen, “A novel error-compensation control for a class of high-order nonlinear systems with input delay,” IEEE transactions on neural networks and learning systems , 2017
2017
Closest in time.
X. Zhao, N. Wang, Y. Zhang, S. Du, Y. Gao, and J. Sun, “Beyond pairwise matching: Person reidentification via high-order relevance learning,” IEEE transactions on neural networks and learning systems , 2017
2017
Closest in time.