Fetching the paper…
Reading the bibliography…
In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel sentence descriptions to explain the content of images.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1988
Earlier work this paper cites.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
A survey of smoothing techniques for me models
S. F. Chen and R. Rosenfeld · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Matching words and pictures
K. Barnard, P. Duygulu, D. Forsyth, N. De Freitas, D. M. Blei, and M. I. Jordan · 2003
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
C.-Y. Lin and F. J. Och · 2004
Earlier work this paper cites.
The iapr tc-12 benchmark: A new evaluation resource for visual information systems
M. Grubinger, P. Clough, H. Müller, and T. Deselaers · 2006
Earlier work this paper cites.
Three new graphical models for statistical language modelling
A. Mnih and G. Hinton · 2007
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Multiple instance metric learning from automatically labeled bags of faces
M. Guillaumin, J. Verbeek, and C. Schmid · 2010
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Cited alongside, same era.
Collecting image annotations using amazon’s mechanical turk
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier · 2010
Cited alongside, same era.
Learning cross-modality similarity for multinomial data
Y. Jia, M. Salzmann, and T. Darrell · 2011
Cited alongside, same era.
Baby talk: Understanding and generating image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Cited alongside, same era.
Extensions of recurrent neural network language model
T. Mikolov, S. Kombrink, L. Burget, J. Cernocky, and S. Khudanpur · 2011
Cited alongside, same era.
From image annotation to image description
A. Gupta and P. Mannem · 2012
Cited alongside, same era.
Multimodal learning with deep boltzmann machines
N. Srivastava and R. Salakhutdinov · 2012
Later among the works it cites.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2013
Later among the works it cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Later among the works it cites.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Later among the works it cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Choosing linguistics over vision to describe images
A. Gupta, Y. Verma, and C. Jawahar · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Efficient backprop
Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 2012
Cited alongside, same era.
Midge: Generating image descriptions from computer vision detections
M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. Berg, K. Yamaguchi, T. Berg, K. Stratos, and H. Daumé III · 2012
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Y. A. L. M. Hodosh and J. Hockenmaier
Cited in the paper.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Later among the works it cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Closest in time.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and L. Fei-Fei · 2014
Closest in time.
Multimodal neural language models
R. Kiros, R. Zemel, and R. Salakhutdinov · 2014
Closest in time.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, Q. Le, C. Manning, and A. Ng · 2014
Closest in time.