Fetching the paper…
Reading the bibliography…
We propose a novel attention based deep learning architecture for visual question answering task (VQA).
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Verbs semantics and lexical selection
Z. Wu and M. Palmer · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Efficient backprop
Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus · 2012
Earlier work this paper cites.
Adadelta: An adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Learning the visual interpretation of sentences
C. L. Zitnick, D. Parikh, and L. Vanderwende · 2013
Earlier work this paper cites.
Multiple object recognition with visual attention
J. Ba, V. Mnih, and K. Kavukcuoglu · 2014
Earlier work this paper cites.
Return of the devil in the details: Delving deep into convolutional nets
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman · 2014
Earlier work this paper cites.
Semantic image segmentation with deep convolutional nets and fully connected crfs
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Later among the works it cites.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko · 2014
Later among the works it cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Are you talking to a machine? dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Long, E. Shelhamer, and T. Darrell · 2014
Cited alongside, same era.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, and A. Yuille · 2014
Cited alongside, same era.
Recurrent models of visual attention
V. Mnih, N. Heess, A. Graves, et al · 2014
Cited alongside, same era.
Attention for fine-grained categorization
P. Sermanet, A. Frome, and E. Real · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
J. Jin, K. Fu, R. Cui, F. Sha, and C. Zhang · 2015
Closest in time.
A dynamic convolutional layer for short range weather prediction
B. Klein, L. Wolf, and Y. Afek · 2015
Closest in time.
Bilinear cnn models for fine-grained visual recognition
T.-Y. Lin, A. RoyChowdhury, and S. Maji · 2015
Closest in time.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Closest in time.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Closest in time.