Fetching the paper…
Reading the bibliography…
Automatically describing an image with a sentence is a long-standing challenge in computer vision and natural language processing.
A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics
D. Martin, C. Fowlkes, D. Tal, and J. Malik · 2001
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
A taxonomy and evaluation of dense two-frame stereo correspondence algorithms
D. Scharstein and R. Szeliski · 2002
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
A. Nenkova and R. J. Passonneau · 2004
Earlier work this paper cites.
Understanding inverse document frequency: On theoretical arguments for idf
S. Robertson · 2004
Earlier work this paper cites.
Rouge: a package for automatic evaluation of summaries
C. yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
The physics of optimal decision making: a formal analysis of models of performance in two-alternative forced-choice tasks
R. Bogacz, E. Brown, J. Moehlis, P. Holmes, and J. D. Cohen · 2006
Earlier work this paper cites.
Re-evaluating the role of bleu in machine translation research
C. Callison-burch and M. Osborne · 2006
Earlier work this paper cites.
Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers
A. Gupta and L. S. Davis · 2008
Earlier work this paper cites.
Utility data annotation with amazon mechanical turk
A. Sorokin and D. Forsyth · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning to detect unseen object classes by betweenclass attribute transfer
C. H. Lampert, H. Nickisch, and S. Harmeling · 2009
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2010 (VOC2010) Results
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Object detection with discriminatively trained part based models
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan · 2010
Earlier work this paper cites.
ImageCLEF: Experimental Evaluation in Visual Information Retrieval
H. Mller, P. Clough, T. Deselaers, and B. Caputo · 2010
Cited alongside, same era.
Collecting image annotations using amazon’s mechanical turk
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier · 2010
Cited alongside, same era.
Baby talk: Understanding and generating image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Cited alongside, same era.
Composing simple image descriptions using web-scale n-grams
S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi · 2011
Cited alongside, same era.
Action recognition from a distributed representation of pose and appearance
S. Maji, L. Bourdev, and J. Malik · 2011
Cited alongside, same era.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Later among the works it cites.
Bringing semantics into focus using visual abstraction
C. L. Zitnick and D. Parikh · 2013
Later among the works it cites.
Learning a recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2014
Closest in time.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Closest in time.
Learning to rank using high-order information
P. K. Dokania, A. Behl, C. V. Jawahar, and P. M. Kumar · 2014
Closest in time.
Long-term recurrent convolutional networks for visual recognition and description
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Relative Attributes
D. Parikh and K. Grauman · 2011
Cited alongside, same era.
Recognition using visual phrases
M. A. Sadeghi and A. Farhadi · 2011
Cited alongside, same era.
Adaptively learning the crowd kernel
O. Tamuz, C. Liu, S. Belongie, O. Shamir, and A. T. Kalai · 2011
Cited alongside, same era.
Corpus-guided sentence generation of natural images
Y. Yang, C. L. Teo, H. D. III, and Y. Aloimonos · 2011
Cited alongside, same era.
Understanding and predicting importance in images
A. C. Berg, T. L. Berg, H. D. III, J. Dodge, A. Goyal, X. Han, A. Mensch, M. Mitchell, A. Sood, K. Stratos, and K. Yamaguchi · 2012
Cited alongside, same era.
Choosing linguistics over vision to describe images
A. Gupta, Y. Verma, and C. Jawahar · 2012
Cited alongside, same era.
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Closest in time.
Comparing automatic evaluation measures for image description
D. Elliott and F. Keller · 2014
Closest in time.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2014
Closest in time.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and L. Fei-Fei · 2014
Closest in time.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Closest in time.
Microsoft COCO: Common objects in context
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Closest in time.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Closest in time.
See no evil, say no evil: Description generation from densely labeled images
M. Yatskar, M. Galley, L. Vanderwende, and L. Zettlemoyer · 2014
Closest in time.
Microsoft COCO Captions: Data Collection and Evaluation Server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick · 2015
Closest in time.