Fetching the paper…
Reading the bibliography…
Visual question answering is fundamentally compositional in nature---a question like "where is the dog?" shares substructure with questions like "what color is the dog?" and "where is the cat?" This paper seeks to simultaneously exploit the representational capacity of deep networks and the compositional linguistic structure of questions.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Accurate unlexicalized parsing
D. Klein and C. D. Manning · 2003
Earlier work this paper cites.
The Stanford typed dependencies representation
M.-C. De Marneffe and C. D. Manning · 2008
Earlier work this paper cites.
A joint model of language and perception for grounded attribute learning
C. Matuszek, N. Fitzgerald, L. Zettlemoyer, L. Bo, and D. Fox · 2012
Earlier work this paper cites.
ADADELTA: An adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Learning distributions over logical forms for referring expression generation
N. FitzGerald, Y. Artzi, and L. Zettlemoyer · 2013
Earlier work this paper cites.
Jointly learning to parse and perceive: connecting natural language to the physical world
J. Krishnamurthy and T. Kollar · 2013
Earlier work this paper cites.
Learning dependency-based compositional semantics
P. Liang, M. I. Jordan, and D. Klein · 2013
Earlier work this paper cites.
Parsing with compositional vector grammars
R. Socher, J. Bauer, C. D. Manning, and A. Y. Ng · 2013
Earlier work this paper cites.
Grounding language with points and paths in continuous spaces
J. Andreas and D. Klein · 2014
Cited alongside, same era.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho, B. van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Cited alongside, same era.
A neural network for factoid question answering over paragraphs
M. Iyyer, J. Boyd-Graber, L. Claudino, R. Socher, and H. Daumé III · 2014
Cited alongside, same era.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and L. Fei-Fei · 2014
Cited alongside, same era.
What are you talking about? text-to-image coreference
Are you talking to a machine? dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Closest in time.
Visual turing test for computer vision systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Closest in time.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Closest in time.
Learning to answer questions from image using convolutional neural network
L. Ma and Z. L. andiyyer Hang Li · 2015
Closest in time.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Closest in time.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
J. Weston, S. Chopra, and A. Bordes · 2014
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
B. Plummer, L. Wang, C. Cervantes, J. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Closest in time.
Image question answering: A visual semantic embedding model and a new dataset
M. Ren, R. Kiros, and R. Zemel · 2015
Closest in time.
Image question answering: A visual semantic embedding model and a new dataset
M. Ren, R. Kiros, and R. S. Zemel · 2015
Closest in time.
Towards ai-complete question answering: a set of prerequisite toy tasks
J. Weston, A. Bordes, S. Chopra, and T. Mikolov · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Closest in time.
Visual madlibs: Fill in the blank image generation and question answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Closest in time.