Fetching the paper…
Reading the bibliography…
The Flickr30k dataset has become a standard benchmark for sentence-based image description.
Relations between two sets of variates
Hotelling, H. (1936) · 1936
Earlier work this paper cites.
Using decision trees for coreference resolution
McCarthy, J. F. and Lehnert, W. G. (1995) · 1995
Earlier work this paper cites.
A machine learning approach to coreference resolution of noun phrases
Soon, W. M., Ng, H. T., and Lim, D. C. Y. (2001) · 2001
Earlier work this paper cites.
The iapr tc-12 benchmark: A new evaluation resource for visual information systems
Grubinger, M., Clough, P., Müller, H., and Deselaers, T. (2006) · 2006
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2008 (VOC2008) Results
Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., and Zisserman, A. (2008) · 2008
Earlier work this paper cites.
Utility data annotation with Amazon Mechanical Turk
Sorokin, A. and Forsyth, D. (2008) · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Farhadi, A., Hejrati, S., Sadeghi, A., Young, P., Rashtchian, C., Hockenmaier, J., and Forsyth, D. A. (2010) · 2010
Earlier work this paper cites.
Cross-caption coreference resolution for automatic image understanding
Hodosh, M., Young, P., Rashtchian, C., and Hockenmaier, J. (2010) · 2010
Earlier work this paper cites.
Improving the Fisher kernel for large-scale image classification
Perronnin, F., Sánchez, J., and Mensink, T. (2010) · 2010
Earlier work this paper cites.
Collecting image annotations using Amazon’s mechanical turk
Rashtchian, C., Young, P., Hodosh, M., and Hockenmaier, J. (2010) · 2010
Earlier work this paper cites.
I2T: Image parsing to text description
Yao, B., Yang, X., Lin, L., Lee, M. W., and Zhu, S.-C. (2010) · 2010
Earlier work this paper cites.
Baby talk: Understanding and generating image descriptions
Kulkarni, G., Premraj, V., Dhar, S., Li, S., Choi, Y., Berg, A. C., and Berg, T. L. (2011) · 2011
Earlier work this paper cites.
Im2Text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T. L. (2011) · 2011
Earlier work this paper cites.
Detecting visual text
Dodge, J., Goyal, A., Han, X., Mensch, A., Mitchell, M., Stratos, K., Yamaguchi, K., Choi, Y., III, H. D., Berg, A. C., and Berg, T. L. (2012) · 2012
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results
Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., and Zisserman, A. (2012) · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, P. K. and Fergus, R. (2012) · 2012
Earlier work this paper cites.
Crowdsourcing annotations for visual object detection
Su, H., Deng, J., and Fei-Fei, L. (2012) · 2012
Earlier work this paper cites.
A sentence is worth a thousand pixels
Fidler, S., Sharma, A., and Urtasun, R. (2013) · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Hodosh, M., Young, P., and Hockenmaier, J. (2013) · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013) · 2013
Cited alongside, same era.
Selective search for object recognition
Uijlings, J., van de Sande, K., Gevers, T., and Smeulders, A. (2013) · 2013
Cited alongside, same era.
Bringing semantics into focus using visual abstraction
Zitnick, C. L. and Parikh, D. (2013) · 2013
Cited alongside, same era.
Deep fragment embeddings for bidirectional image sentence mapping
Karpathy, A., Joulin, A., and Fei-Fei, L. (2014) · 2014
Cited alongside, same era.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S., Ordonez, V., Matten, M., and Berg, T. (2014) · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
Fast r-cnn
Girshick, R. (2015) · 2015
Closest in time.
Image retrieval using scene graphs
Johnson, J., Krishna, R., Stark, M., Li, L.-J., Shamma, D. A., Bernstein, M., and Fei-Fei, L. (2015) · 2015
Closest in time.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L. (2015) · 2015
Closest in time.
Phrase-based image captioning
Lebret, R., Pinheiro, P. O., and Collobert, R. (2015) · 2015
Closest in time.
Multimodal convolutional neural networks for matching image and sentence
Ma, L., Lu, Z., Shang, L., and Li, H. (2015) · 2015
Closest in time.
Deep captioning with multimodal recurrent neural networks (m-RNN)
Mao, J., Xu, W., Yang, Y., Wang, J., and Yuille, A. (2015) · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kiros, R., Salakhutdinov, R., and Zemel, R. S. (2014) · 2014
Cited alongside, same era.
Fisher vectors derived from hybrid gaussian-laplacian mixture models for image annotation
Klein, B., Lev, G., Sadeh, G., and Wolf, L. (2014) · 2014
Cited alongside, same era.
What are you talking about? text-to-image coreference
Kong, C., Lin, D., Bansal, M., Urtasun, R., and Fidler, S. (2014) · 2014
Cited alongside, same era.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. (2014) · 2014
Cited alongside, same era.
A multi-world approach to question answering about real-world scenes based on uncertain input
Malinowski, M. and Fritz, M. (2014) · 2014
Cited alongside, same era.
Linking people in videos with ”their” names using coreference resolution
Ramanathan, V., Joulin, A., Liang, P., and Fei-Fei, L. (2014) · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S. (2015) · 2015
Closest in time.
Exploring models and data for image question answering
Ren, M., Kiros, R., and Zemel, R. (2015) · 2015
Closest in time.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015) · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Courville, A., Salakhutdinov, R., Zemel, R., and Bengio, Y. (2015) · 2015
Closest in time.
Visual Madlibs: Fill in the blank Image Generation and Question Answering
Yu, L., Park, E., Berg, A. C., and Berg, T. L. (2015) · 2015
Closest in time.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A., Park, D. H., Yang, D., Rohrbach, A., Darrell, T., and Rohrbach, M. (2016) · 2016
Closest in time.
Natural language object retrieval
Hu, R., Xu, H., Rohrbach, M., Feng, J., Saenko, K., and Darrell, T. (2016) · 2016
Closest in time.
Densecap: Fully convolutional localization networks for dense captioning
Johnson, J., Karpathy, A., and Fei-Fei, L. (2016) · 2016
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., Bernstein, M., and Fei-Fei, L. (2016) · 2016
Closest in time.
RNN fisher vectors for action recognition and image annotation
Lev, G., Sadeh, G., Klein, B., and Wolf, L. (2016) · 2016
Closest in time.
Generation and comprehension of unambiguous object descriptions
Mao, J., Jonathan, H., Toshev, A., Camburu, O., Yuille, A., and Murphy, K. (2016) · 2016
Closest in time.
Grounding of textual phrases in images by reconstruction
Rohrbach, A., Rohrbach, M., Hu, R., Darrell, T., and Schiele, B. (2016) · 2016
Closest in time.
Solving visual madlibs with multiple cues
Tommasi, T., Mallya, A., Plummer, B. A., Lazebnik, S., Berg, A., and Berg., T. (2016) · 2016
Closest in time.
Top-down neural attention by excitation backprop
Zhang, J., Lin, Z., Brandt, Jonathan, S. X., and Sclaroff, S. (2016) · 2016
Closest in time.