Fetching the paper…
Reading the bibliography…
We propose the task of free-form and open-ended Visual Question Answering (VQA).
Building Large Knowledge-Based Systems; Representation and Inference in the Cyc Project
D. B. Lenat and R. V. Guha · 1989
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
K. Toutanova, D. Klein, C. D. Manning, and Y. Singer · 2003
Earlier work this paper cites.
ConceptNet — A Practical Commonsense Reasoning Tool-Kit
H. Liu and P. Singh · 2004
Earlier work this paper cites.
Freebase: A Collaboratively Created Graph Database for Structuring Human Knowledge
K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor · 2008
Earlier work this paper cites.
VizWiz: Nearly Real-time Answers to Visual Questions
J. P. Bigham, C. Jayant, H. Ji, G. Little, A. Miller, R. C. Miller, R. Miller, A. Tatarowicz, B. White, S. White, and T. Yeh · 2010
Earlier work this paper cites.
Toward an Architecture for Never-Ending Language Learning
A. Carlson, J. Betteridge, B. Kisiel, B. Settles, E. R. H. Jr., and T. M. Mitchell · 2010
Earlier work this paper cites.
Every Picture Tells a Story: Generating Sentences for Images
A. Farhadi, M. Hejrati, A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Hierarchical Semantic Indexing for Large Scale Image Retrieval
J. Deng, A. C. Berg, and L. Fei-Fei · 2011
Earlier work this paper cites.
Baby Talk: Understanding and Generating Simple Image Descriptions
G. Kulkarni, V. Premraj, S. L. Sagnik Dhar and, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Midge: Generating Image Descriptions From Computer Vision Detections
M. Mitchell, J. Dodge, A. Goyal, K. Yamaguchi, K. Stratos, X. Han, A. Mensch, A. Berg, T. L. Berg, and H. Daume III · 2012
Earlier work this paper cites.
NEIL: Extracting Visual Knowledge from Web Data
X. Chen, A. Shrivastava, and A. Gupta · 2013
Earlier work this paper cites.
Paraphrase-Driven Learning for Open Question Answering
A. Fader, L. Zettlemoyer, and O. Etzioni · 2013
Earlier work this paper cites.
Reporting bias and knowledge extraction
J. Gordon and B. V. Durme · 2013
Earlier work this paper cites.
YouTube2Text: Recognizing and Describing Arbitrary Activities Using Semantic Hierarchies and Zero-Shot Recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko · 2013
Earlier work this paper cites.
Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Earlier work this paper cites.
Distributed Representations of Words and Phrases and their Compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Attributes in visual reference
M. Mitchell, K. van Deemter, and E. Reiter · 2013
Earlier work this paper cites.
Generating Expressions that Refer to Visible Objects
M. Mitchell, K. Van Deemter, and E. Reiter · 2013
Earlier work this paper cites.
MCTest: A Challenge Dataset for the Open-Domain Machine Comprehension of Text
M. Richardson, C. J. Burges, and E. Renshaw · 2013
Earlier work this paper cites.
Translating Video Content to Natural Language Descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Cited alongside, same era.
Bringing Semantics Into Focus Using Visual Abstraction
C. L. Zitnick and D. Parikh · 2013
Cited alongside, same era.
Learning the Visual Interpretation of Sentences
C. L. Zitnick, D. Parikh, and L. Vanderwende · 2013
Cited alongside, same era.
Zero-Shot Learning via Visual Abstraction
S. Antol, C. L. Zitnick, and D. Parikh · 2014
Cited alongside, same era.
Dynamic wordclouds and vennclouds for exploratory data analysis
G. Coppersmith and E. Kelly · 2014
Cited alongside, same era.
Comparing Automatic Evaluation Measures for Image Description
D. Elliott and F. Keller · 2014
Cited alongside, same era.
Mind’s Eye: A Recurrent Visual Representation for Image Caption Generation
X. Chen and C. L. Zitnick · 2015
Closest in time.
Long-term Recurrent Convolutional Networks for Visual Recognition and Description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Closest in time.
From Captions to Visual Concepts and Back
H. Fang, S. Gupta, F. N. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2015
Closest in time.
Are you talking to a machine? dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, and A. Yuille · 2015
Closest in time.
Deep Visual-Semantic Alignments for Generating Image Descriptions
A. Karpathy and L. Fei-Fei · 2015
Closest in time.
Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Open Question Answering over Curated and Extracted Knowledge Bases
A. Fader, L. Zettlemoyer, and O. Etzioni · 2014
Cited alongside, same era.
A Visual Turing Test for Computer Vision Systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
ReferItGame: Referring to Objects in Photographs of Natural Scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Cited alongside, same era.
What Are You Talking About? Text-to-Image Coreference
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Cited alongside, same era.
Microsoft COCO: Common Objects in Context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2015
Closest in time.
R. Kiros, Y. Zhu, R. Salakhutdinov, R. S. Zemel, A. Torralba, R. Urtasun, and S. Fidler · 2015
Closest in time.
Don’t Just Listen, Use Your Imagination: Leveraging Visual Common Sense for Non-Visual Tasks
X. Lin and D. Parikh · 2015
Closest in time.
Don’t just listen, use your imagination: Leveraging visual common sense for non-visual tasks
X. Lin and D. Parikh · 2015
Closest in time.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Closest in time.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Closest in time.
Viske: Visual knowledge extraction and question answering by visual verification of relation phrases
F. Sadeghi, S. K. Kumar Divvala, and A. Farhadi · 2015
Closest in time.
CIDEr: Consensus-based Image Description Evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Learning common sense through visual abstraction
R. Vendantam, X. Lin, T. Batra, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Show and Tell: A Neural Image Caption Generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.
Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
J. Weston, A. Bordes, S. Chopra, and T. Mikolov · 2015
Closest in time.
Visual madlibs: Fill-in-the-blank description generation and question answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Closest in time.
Yin and yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2015
Closest in time.
Adopting Abstract Images for Semantic Scene Understanding
C. L. Zitnick, R. Vedantam, and D. Parikh · 2015
Closest in time.