Fetching the paper…
Reading the bibliography…
We present a model that generates natural language descriptions of images and their regions.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K. K. Paliwal · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Matching words and pictures
K. Barnard, P. Duygulu, D. Forsyth, N. De Freitas, D. M. Blei, and M. I. Jordan · 2003
Earlier work this paper cites.
Neural probabilistic language models
Y. Bengio, H. Schwenk, J.-S. Senécal, F. Morin, and J.-L. Gauvain · 2006
Earlier work this paper cites.
What do we perceive in a glance of a real-world scene?
L. Fei-Fei, A. Iyer, C. Koch, and P. Perona · 2007
Earlier work this paper cites.
What, where and who? classifying events by scene and object recognition
L.-J. Li and L. Fei-Fei · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Decomposing a scene into geometric and semantically consistent regions
S. Gould, R. Fulton, and D. Koller · 2009
Earlier work this paper cites.
Towards total scene understanding: Classification, annotation and segmentation in an automatic framework
L.-J. Li, R. Socher, and L. Fei-Fei · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora
R. Socher and L. Fei-Fei · 2010
Earlier work this paper cites.
I2t: Image parsing to text description
B. Z. Yao, X. Yang, L. Lin, M. W. Lee, and S.-C. Zhu · 2010
Earlier work this paper cites.
Learning cross-modality similarity for multinomial data
Y. Jia, M. Salzmann, and T. Darrell · 2011
Earlier work this paper cites.
Baby talk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Cited alongside, same era.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. E. Hinton · 2011
Cited alongside, same era.
Corpus-guided sentence generation of natural images
Y. Yang, C. L. Teo, H. Daumé III, and Y. Aloimonos · 2011
Cited alongside, same era.
A. Barbu, A. Bridge, Z. Burchill, D. Coroian, S. Dickinson, S. Fidler, A. Michaux, S. Mussman, S. Narayanaswamy, D. Salvi, et al · 2012
Cited alongside, same era.
From image annotation to image description
A. Gupta and P. Mannem · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Closest in time.
Glove: Global vectors for word representation
R. JeffreyPennington and C. Manning · 2014
Closest in time.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and L. Fei-Fei · 2014
Closest in time.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Closest in time.
Multimodal neural language models
R. Kiros, R. S. Zemel, and R. Salakhutdinov · 2014
Closest in time.
What are you talking about? text-to-image coreference
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Collective generation of natural image descriptions
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi · 2012
Cited alongside, same era.
A Joint Model of Language and Perception for Grounded Attribute Learning
C. Matuszek*, N. FitzGerald*, L. Zettlemoyer, L. Bo, and D. Fox · 2012
Cited alongside, same era.
Midge: Generating image descriptions from computer vision detections
M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. Berg, K. Yamaguchi, T. Berg, K. Stratos, and H. Daumé, III · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude., 2012
T. Tieleman and G. E. Hinton · 2012
Cited alongside, same era.
Image description using visual dependency representations
D. Elliott and F. Keller · 2013
Cited alongside, same era.
A sentence is worth a thousand pixels
S. Fidler, A. Sharma, and R. Urtasun · 2013
Cited alongside, same era.
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Closest in time.
Treetalk: Composition and compression of trees for image descriptions
P. Kuznetsova, V. Ordonez, T. L. Berg, U. C. Hill, and Y. Choi · 2014
Closest in time.
Visual semantic search: Retrieving videos via complex textual queries
D. Lin, S. Fidler, C. Kong, and R. Urtasun · 2014
Closest in time.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Closest in time.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Closest in time.
Imagenet large scale visual recognition challenge, 2014
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2014
Closest in time.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Closest in time.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng · 2014
Closest in time.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Closest in time.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2014
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Closest in time.
See no evil, say no evil: Description generation from densely labeled images
M. Yatskar, L. Vanderwende, and L. Zettlemoyer · 2014
Closest in time.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Closest in time.
Recurrent neural network regularization
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Closest in time.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick · 2015
Closest in time.