Fetching the paper…
Reading the bibliography…
This work aims to address the problem of image-based question-answering (QA) with new models and datasets.
N. Chomsky, Conditions on Transformations . New York: Academic Press, 1973
1973
Earlier work this paper cites.
Z. Wu and M. Palmer, “Verb semantics and lexical selection,” in ACL , 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
C. Fellbaum, Ed., WordNet An Electronic Lexical Database . Cambridge, MA; London: The MIT Press, May 1998
1998
Earlier work this paper cites.
D. Klein and C. D. Manning, “Accurate unlexicalized parsing,” in ACL , 2003
2003
Earlier work this paper cites.
S. Bird, “NLTK: the natural language toolkit,” in ACL , 2006
2006
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. L. Berg, “Im2text: Describing images using 1 million captioned photographs,” in NIPS , 2011
2011
Earlier work this paper cites.
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “Indoor segmentation and support inference from RGBD images,” in ECCV , 2012
2012
Earlier work this paper cites.
A. Karpathy, A. Joulin, and L. Fei-Fei, “Deep fragment embeddings for bidirectional image sentence mapping,” in NIPS , 2013
2013
Earlier work this paper cites.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” in ICLR , 2013
2013
Earlier work this paper cites.
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov, “DeViSE: A deep visual-semantic embedding model,” in NIPS , 2013
2013
Earlier work this paper cites.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,” J. Artif. Intell. Res. (JAIR) , vol. 47, pp. 853–899, 2013
2013
Cited alongside, same era.
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille, “Explain images with multimodal recurrent neural networks,” NIPS Deep Learning Workshop , 2014
2014
Cited alongside, same era.
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR , 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
M. Malinowski and M. Fritz, “Towards a visual Turing challenge,” in NIPS Workshop on Learning Semantics , 2014
R. Lebret, P. O. Pinheiro, and R. Collobert, “Phrase-based image captioning,” in ICML , 2015
2015
Closest in time.
B. Klein, G. Lev, G. Lev, and L. Wolf, “Fisher vectors derived from hybrid Gaussian-Laplacian mixture models for image annotations,” in CVPR , 2015
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common Objects in Context,” in ECCV , 2014
2014
Cited alongside, same era.
M. Malinowski and M. Fritz, “A multi-world approach to question answering about real-world scenes based on uncertain input,” in NIPS , 2014
2014
Cited alongside, same era.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in CVPR , 2015
2015
Cited alongside, same era.
R. Kiros, R. Salakhutdinov, and R. S. Zemel, “Unifying visual-semantic embeddings with multimodal neural language models,” TACL , 2015
2015
Cited alongside, same era.
H. Fang, S. Gupta, F. N. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig, “From captions to visual concepts and back,” in CVPR , 2015
2015
Cited alongside, same era.
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in ICML , 2015
2015
Cited alongside, same era.
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR , 2015
2015
Closest in time.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and L. Fei-Fei, “Imagenet large scale visual recognition challenge,” IJCV , 2015
2015
Closest in time.