Fetching the paper…
Reading the bibliography…
We propose a method for visual question answering which combines an internal representation of the content of an image with information extracted from a general knowledge base to answer a broad range of image-based questions.
Verbs semantics and lexical selection
Z. Wu and M. Palmer · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
S. Auer, C. Bizer, G. Kobilarov, J. Lehmann, R. Cyganiak, and Z. Ives · 2007
Earlier work this paper cites.
Freebase: a collaboratively created graph database for structuring human knowledge
K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Building Watson: An overview of the DeepQA project
D. Ferrucci, E. Brown, J. Chu-Carroll, J. Fan, D. Gondek, A. A. Kalyanpur, A. Lally, J. W. Murdock, E. Nyberg, J. Prager, et al · 2010
Earlier work this paper cites.
Semantic Parsing on Freebase from Question-Answer Pairs
J. Berant, A. Chou, R. Frostig, and P. Liang · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, C. Gulcehre, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and F. F. Li · 2014
Earlier work this paper cites.
Distributed representations of sentences and documents
Q. V. Le and T. Mikolov · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Towards a Visual Turing Challenge
M. Malinowski and M. Fritz · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Joint video and text parsing for understanding events and answering queries
K. Tu, M. Meng, M. W. Lee, T. E. Choe, and S.-C. Zhu · 2014
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Cited alongside, same era.
Visual Turing test for computer vision systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Closest in time.
Don’t Just Listen, Use Your Imagination: Leveraging Visual Common Sense for Non-Visual Tasks
X. Lin and D. Parikh · 2015
Closest in time.
Learning to Answer Questions From Image using Convolutional Neural Network
L. Ma, Z. Lu, and H. Li · 2015
Closest in time.
Ask Your Neurons: A Neural-based Approach to Answering Questions about Images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Closest in time.
Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
J. Mao, W. Xu, Y. Yang, J. Wang, and A. Yuille · 2015
Closest in time.
Multiscale Combinatorial Grouping for Image Segmentation and Object Proposal Generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Wei, W. Xia, J. Huang, B. Ni, J. Dong, Y. Zhao, and S. Yan · 2014
Cited alongside, same era.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Microsoft COCO captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick · 2015
Cited alongside, same era.
Learning a Recurrent Visual Representation for Image Caption Generation
X. Chen and C. L. Zitnick · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Cited alongside, same era.
J. Pont-Tuset, P. Arbeláez, J. Barron, F. Marques, and J. Malik · 2015
Closest in time.
Image Question Answering: A Visual Semantic Embedding Model and a New Dataset
M. Ren, R. Kiros, and R. Zemel · 2015
Closest in time.
VisKE: Visual Knowledge Extraction and Question Answering by Visual Verification of Relation Phrases
F. Sadeghi, S. K. Kumar Divvala, and A. Farhadi · 2015
Closest in time.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Closest in time.
Building a Large-scale Multimodal Knowledge Base for Visual Question Answering
Y. Zhu, C. Zhang, C. Ré, and L. Fei-Fei · 2015
Closest in time.
What Value Do Explicit High Level Concepts Have in Vision to Language Problems?
Q. Wu, C. Shen, A. v. d. Hengel, L. Liu, and A. Dick · 2016
Closest in time.