Fetching the paper…
Reading the bibliography…
Problems at the intersection of vision and language are of significant importance both as challenging research questions and for the rich set of applications they enable.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Unbiased look at dataset bias
A. Torralba and A. Efros · 2011
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Explain Images with Multimodal Recurrent Neural Networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Mind’s Eye: A Recurrent Visual Representation for Image Caption Generation
X. Chen and C. L. Zitnick · 2015
Earlier work this paper cites.
Exploring nearest neighbor approaches for image captioning
J. Devlin, S. Gupta, R. B. Girshick, M. Mitchell, and C. L. Zitnick · 2015
Earlier work this paper cites.
Long-term Recurrent Convolutional Networks for Visual Recognition and Description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
From Captions to Visual Concepts and Back
H. Fang, S. Gupta, F. N. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, and A. Yuille · 2015
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2015
Earlier work this paper cites.
Deeper LSTM and normalized CNN Visual Question Answering model
J. Lu, X. Lin, D. Batra, and D. Parikh · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Show and Tell: A Neural Image Caption Generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Explicit knowledge-based reasoning for visual question answering
P. Wang, Q. Wu, C. Shen, A. van den Hengel, and A. R. Dick · 2015
Cited alongside, same era.
Visual Madlibs: Fill-in-the-blank Description Generation and Question Answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Cited alongside, same era.
Learning Deep Features for Discriminative Localization
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2015
Cited alongside, same era.
Simple Baseline for Visual Question Answering
B. Zhou, Y. Tian, S. Sukhbaatar, A. Szlam, and R. Fergus · 2015
Cited alongside, same era.
Multimodal Residual Learning for Visual QA
J.-H. Kim, S.-W. Lee, D.-H. Kwak, M.-O. Heo, J. Kim, J.-W. Ha, and B.-T. Zhang · 2016
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2016
Closest in time.
Hierarchical Question-Image Co-Attention for Visual Question Answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Closest in time.
Training recurrent answering units with joint loss minimization for vqa
H. Noh and B. Han · 2016
Closest in time.
Question Relevance in VQA: Identifying Non-Visual And False-Premise Questions
A. Ray, G. Christie, M. Bansal, D. Batra, and D. Parikh · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Analyzing the Behavior of Visual Question Answering Models
A. Agrawal, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Deep compositional question answering with neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Towards Transparent AI Systems: Interpreting Visual Question Answering Models
Y. Goyal, A. Mohapatra, D. Parikh, and D. Batra · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Generating visual explanations
L. A. Hendricks, Z. Akata, M. Rohrbach, J. Donahue, B. Schiele, and T. Darrell · 2016
Cited alongside, same era.
Focused evaluation for image description with binary forced-choice tasks
M. Hodosh and J. Hockenmaier · 2016
Cited alongside, same era.
”Why Should I Trust You?”: Explaining the Predictions of Any Classifier
M. T. Ribeiro, S. Singh, and C. Guestrin · 2016
Closest in time.
Dualnet: Domain-invariant network for visual question answering
K. Saito, A. Shin, Y. Ushiku, and T. Harada · 2016
Closest in time.
R. R. Selvaraju, A. Das, R. Vedantam, M. Cogswell, D. Parikh, and D. Batra · 2016
Closest in time.
Where to look: Focus regions for visual question answering
K. J. Shih, S. Singh, and D. Hoiem · 2016
Closest in time.
The Color of the Cat is Gray: 1 Million Full-Sentences Visual Question Answering (FSVQA)
A. Shin, Y. Ushiku, and T. Harada · 2016
Closest in time.
MovieQA: Understanding Stories in Movies through Question-Answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Closest in time.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Q. Wu, P. Wang, C. Shen, A. van den Hengel, and A. R. Dick · 2016
Closest in time.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Closest in time.
Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering
H. Xu and K. Saenko · 2016
Closest in time.
Stacked Attention Networks for Image Question Answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Closest in time.
Yin and Yang: Balancing and Answering Binary Visual Questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Closest in time.