The SUN attribute database: Beyond categories for deeper scene understanding
G. Patterson, C. Xu, H. Su, and J. Hays · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. Le · 2014
Later among the works it cites.
Joint video and text parsing for understanding events and answering queries
K. Tu, M. Meng, M. W. Lee, T. E. Choe, and S.-C. Zhu · 2014
Later among the works it cites.
Visualizing and understanding convolutional networks
M. Zeiler and R. Fergus · 2014
Later among the works it cites.
DimmWitted: A study of main-memory statistical analytics
C. Zhang and C. Ré · 2014
Later among the works it cites.
Learning Deep Features for Scene Recognition using Places Database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Later among the works it cites.
Reasoning about object affordances in a knowledge base representation
Y. Zhu, A. Fathi, and L. Fei-Fei · 2014
Later among the works it cites.
VQA: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, L. Zitnick, and D. Parikh · 2015
Closest in time.
Mining semantic affordances of visual object categories
Y.-W. Chao, Z. Wang, R. Mihalcea, and J. Deng · 2015
Closest in time.
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2015
Closest in time.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Closest in time.
Are you talking to a machine? dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Closest in time.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Closest in time.
Generating multi-sentence lingual descriptions of indoor scenes
D. Lin, C. Kong, S. Fidler, and R. Urtasun · 2015
Closest in time.
Don’t just listen, use your imagination: Leveraging visual common sense for non-visual tasks
X. Lin and D. Parikh · 2015
Closest in time.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Closest in time.
Viske: Visual knowledge extraction and question answering by visual verification of relation phrases
F. Sadeghi, S. K. Divvala, and A. Farhadi · 2015
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Closest in time.
Visual Madlibs: Fill in the blank Image Generation and Question Answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Closest in time.