Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) requires AI models to comprehend data in two domains, vision and text.
Tucker, L. R., 1966. Some Mathematical Notes on Three-Mode Factor Analysis. Psychometrika 31 (3), 279–311
1966
Earlier work this paper cites.
Rensink, R. A., Jan. 2000. The Dynamic Representation of Scenes. Vis. Cogn. 7 (1-3), 17–42
2000
Earlier work this paper cites.
Charikar, M., Chen, K., Farach-Colton, M., 2004. Finding frequent items in data streams. Theor. Comput. Sci. 312 (1), 3–15
2004
Earlier work this paper cites.
Glorot, X., Bengio, Y., 2010. Understanding the Difficulty of Training Deep Feedforward Neural Networks. In: AISTATS. pp. 249–256
2010
Earlier work this paper cites.
2014
Earlier work this paper cites.
Kingma, D. P., Ba, J., 2014. Adam: A Method for Stochastic Optimization. arXiv:1412.6980
2014
Earlier work this paper cites.
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C. L., 2014. Microsoft Coco: Common Objects in Context. In: ECCV. pp. 740–755
2014
Earlier work this paper cites.
Mnih, V., Heess, N., Graves, A., Kavukcuoglu, K., 2014. Recurrent Models of Visual Attention. In: Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2. NIPS’14. MIT Press, Cambridge, MA, USA, pp. 2204–2212
2014
Earlier work this paper cites.
Pennington, J., Socher, R., Manning, C. D., 2014. GloVe: Global Vectors for Word Representation. In: EMNLP. pp. 1532–1543
2014
Earlier work this paper cites.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Zitnick, C. L., Batra, D., Parikh, D., 2015. VQA: Visual Question Answering. In: CVPR. pp. 2425–2433
2015
Earlier work this paper cites.
Bahdanau, D., Cho, K., Bengio, Y., 2015. Neural Machine Translation by Jointly Learning to Align and Translate. In: ICLR
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J., 2015. Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. In: ICCV. pp. 1026–1034
2015
Earlier work this paper cites.
Kalchbrenner, N., Danihelka, I., Graves, A., 2015. Grid long short-term memory. arXiv:1507.01526
2015
Earlier work this paper cites.
Kumar, P. R., Varaiya, P., 2015. Stochastic Systems: Estimation, Identification, and Adaptive Control. SIAM
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Fukui, A., Park, D. H., Yang, D., Rohrbach, A., Darrell, T., Rohrbach, M., 2016. Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding. In: EMNLP. pp. 457–468
2016
Cited alongside, same era.
Gao, Y., Beijbom, O., Zhang, N., Darrell, T., 2016. Compact bilinear pooling. In: CVPR. pp. 317–326
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D., 2017. Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In: CVPR. pp. 6904–6913
2017
Later among the works it cites.
Kim, J.-H., On, K.-W., Kim, J., Ha, J.-W., Zhang, B.-T., 2017. Hadamard Product for Low-Rank Bilinear Pooling. In: ICLR
2017
Later among the works it cites.
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., others, 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis. 123 (1), 32–73
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Lu, J., Yang, J., Batra, D., Parikh, D., 2016. Hierarchical Question-Image Co-Attention for Visual Question Answering. In: NIPS. pp. 289–297
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Xiong, C., Merity, S., Socher, R., 2016. Dynamic Memory Networks for Visual and Textual Question Answering. In: ICML. pp. 2397–2406
2016
Cited alongside, same era.
Xu, H., Saenko, K., 2016. Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering. In: ECCV. pp. 451–466
2016
Cited alongside, same era.
2017
Cited alongside, same era.
Arras, L., Montavon, G., Müller, K.-R., Samek, W., 2017. Explaining Recurrent Neural Network Predictions in Sentiment Analysis. In: EMNLP’17 Workshop on Computational Approaches to Subjectivity, Sentiment & Social Media Analysis (WASSA). pp. 159–168
2017
Cited alongside, same era.
Nam, H., Ha, J.-W., Kim, J., 2017. Dual Attention Networks for Multimodal Reasoning and Matching. In: CVPR. pp. 299–307
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Bosse, S., Maniry, D., Müller, K.-R., Wiegand, T., Samek, W., 2018. Deep Neural Networks for No-Reference and Full-Reference Image Quality Assessment. IEEE Trans. Image Process. 27 (1), 206–219
2018
Closest in time.
Homayounfar, N., Ma, W., Lakshmikanth, S. K., Urtasun, R., Jun. 2018. Hierarchical Recurrent Attention Networks for Structured Online Maps. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3417–3426
2018
Closest in time.
Jiang, Y., Natarajan, V., Chen, X., Rohrbach, M., Batra, D., Parikh, D., Jul. 2018. Pythia v0.1: The Winning Entry to the VQA Challenge 2018. ArXiv180709956 Cs
2018
Closest in time.
Montavon, G., Samek, W., Müller, K.-R., 2018. Methods for Interpreting and Understanding Deep Neural Networks. Digit. Signal Process. 73, 1–15
2018
Closest in time.