Fetching the paper…
Reading the bibliography…
In this paper we describe the Microsoft COCO Caption dataset and evaluation server.
G. A. Miller, “Wordnet: a lexical database for english,” Communications of the ACM , vol. 38, no. 11, pp. 39–41, 1995
1995
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
K. Barnard and D. Forsyth, “Learning the semantics of words and pictures,” in ICCV , vol. 2, 2001, pp. 408–415
2001
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in ACL , 2002
2002
Earlier work this paper cites.
K. Barnard, P. Duygulu, D. Forsyth, N. De Freitas, D. M. Blei, and M. I. Jordan, “Matching words and pictures,” JMLR , vol. 3, pp. 1107–1135, 2003
2003
Earlier work this paper cites.
V. Lavrenko, R. Manmatha, and J. Jeon, “A model for learning the semantics of pictures,” in NIPS , 2003
2003
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in ACL Workshop , 2004
2004
Earlier work this paper cites.
M. Grubinger, P. Clough, H. Müller, and T. Deselaers, “The iapr tc-12 benchmark: A new evaluation resource for visual information systems,” in LREC Workshop on Language Resources for Content-based Image Retrieval , 2006
2006
Earlier work this paper cites.
C. Callison-Burch, M. Osborne, and P. Koehn, “Re-evaluation the role of bleu in machine translation research.” in EACL , vol. 6, 2006, pp. 249–256
2006
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in CVPR , 2009
2009
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth, “Every picture tells a story: Generating sentences from images,” in ECCV , 2010
2010
Earlier work this paper cites.
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg, “Baby talk: Understanding and generating simple image descriptions,” in CVPR , 2011
2011
Earlier work this paper cites.
Y. Yang, C. L. Teo, H. Daumé III, and Y. Aloimonos, “Corpus-guided sentence generation of natural images,” in EMNLP , 2011
2011
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing images using 1 million captioned photographs.” in NIPS , 2011
2011
Earlier work this paper cites.
M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. Berg, K. Yamaguchi, T. Berg, K. Stratos, and H. Daumé III, “Midge: Generating image descriptions from computer vision detections,” in EACL , 2012
2012
Earlier work this paper cites.
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi, “Collective generation of natural image descriptions,” in ACL , 2012
2012
Earlier work this paper cites.
A. Gupta, Y. Verma, and C. Jawahar, “Choosing linguistics over vision to describe images.” in AAAI , 2012
2012
Cited alongside, same era.
E. Bruni, G. Boleda, M. Baroni, and N.-K. Tran, “Distributional semantics in technicolor,” in ACL , 2012
2012
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet classification with deep convolutional neural networks,” in NIPS , 2012
2012
Cited alongside, same era.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics.” JAIR , vol. 47, pp. 853–899, 2013
2013
Cited alongside, same era.
Y. Feng and M. Lapata, “Automatic caption generation for news images,” TPAMI , vol. 35, no. 4, pp. 797–812, 2013
2013
Cited alongside, same era.
2014
Later among the works it cites.
2014
Later among the works it cites.
2014
Later among the works it cites.
2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Elliott and F. Keller, “Image description using visual dependency representations.” in EMNLP , 2013, pp. 1292–1302
2013
Cited alongside, same era.
A. Karpathy, A. Joulin, and F.-F. Li, “Deep fragment embeddings for bidirectional image sentence mapping,” in NIPS , 2014
2014
Cited alongside, same era.
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik, “Improving image-sentence embeddings using large weakly annotated photo collections,” in ECCV , 2014, pp. 529–545
2014
Cited alongside, same era.
R. Mason and E. Charniak, “Nonparametric method for data-driven image captioning,” in ACL , 2014
2014
Cited alongside, same era.
P. Kuznetsova, V. Ordonez, T. Berg, and Y. Choi, “Treetalk: Composition and compression of trees for image descriptions,” TACL , vol. 2, pp. 351–362, 2014
2014
Cited alongside, same era.
K. Ramnath, S. Baker, L. Vanderwende, M. El-Saban, S. N. Sinha, A. Kannan, N. Hassan, M. Galley, Y. Yang, D. Ramanan, A. Bergamo, and L. Torresani, “Autocaption: Automatic caption generation for personal photos,” in WACV , 2014
2014
Cited alongside, same era.
A. Lazaridou, E. Bruni, and M. Baroni, “Is this a wampimuk? cross-modal mapping between distributional semantics and the visual world,” in ACL , 2014
2014
Cited alongside, same era.
2014
Later among the works it cites.
2014
Later among the works it cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” TACL , vol. 2, pp. 67–78, 2014
2014
Later among the works it cites.
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in ECCV , 2014
2014
Later among the works it cites.
M. Denkowski and A. Lavie, “Meteor universal: Language specific translation evaluation for any target language,” in EACL Workshop on Statistical Machine Translation , 2014
2014
Later among the works it cites.
2014
Later among the works it cites.
C. D. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. J. Bethard, and D. McClosky, “The Stanford CoreNLP natural language processing toolkit,” in Proceedings of 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations , 2014, pp. 55–60. [Online]. Available: http://www.aclweb.org/anthology/P/P14/P14-5010
2014
Later among the works it cites.
D. Elliott and F. Keller, “Comparing automatic evaluation measures for image description,” in Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics , vol. 2, 2014, pp. 452–457
2014
Later among the works it cites.
2015
Closest in time.
2015
Closest in time.
J. Chen, P. Kuznetsova, D. Warren, and Y. Choi, “Déjá image-captions: A corpus of expressive image descriptions in repetition,” in NAACL , 2015
2015
Closest in time.