Fetching the paper…
Reading the bibliography…
In this paper we explore the bi-directional mapping between images and their sentence-based descriptions.
Words versus objects: Comparison of free verbal recall
L. R. Lieberman and J. T. Culpepper · 1965
Earlier work this paper cites.
Why are pictures easier to recall than words?
A. Paivio, T. B. Rogers, and P. C. Smythe · 1968
Earlier work this paper cites.
Experimental analysis of the real-time recurrent learning algorithm
R. J. Williams and D. Zipser · 1989
Earlier work this paper cites.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Imagery in sentence comprehension: an fmri study
M. A. Just, S. D. Newman, T. A. Keller, A. McEleney, and P. A. Carpenter · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Neural probabilistic language models
Y. Bengio, H. Schwenk, J.-S. Senécal, F. Morin, and J.-L. Gauvain · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Collecting image annotations using Amazon’s mechanical turk
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier · 2010
Cited alongside, same era.
I2T: Image parsing to text description
B. Z. Yao, X. Yang, L. Lin, M. W. Lee, and S.-C. Zhu · 2010
Cited alongside, same era.
Baby talk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Cited alongside, same era.
Strategies for training large scale neural network language models
T. Mikolov, A. Deoras, D. Povey, L. Burget, and J. Cernocky · 2011
Cited alongside, same era.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. E. Hinton · 2011
Cited alongside, same era.
Corpus-guided sentence generation of natural images
Y. Yang, C. L. Teo, H. Daumé III, and Y. Aloimonos · 2011
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Later among the works it cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Later among the works it cites.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, Q. Le, C. Manning, and A. Ng · 2013
Later among the works it cites.
Bringing semantics into focus using visual abstraction
C. L. Zitnick and D. Parikh · 2013
Later among the works it cites.
Comparing automatic evaluation measures for image description
D. Elliott and F. Keller · 2014
Closest in time.
Rich feature hierarchies for accurate object detection and semantic segmentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Choosing linguistics over vision to describe images
A. Gupta, Y. Verma, and C. Jawahar · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Collective generation of natural image descriptions
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi · 2012
Cited alongside, same era.
Context dependent recurrent neural network language model
T. Mikolov and G. Zweig · 2012
Cited alongside, same era.
Midge: Generating image descriptions from computer vision detections
M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. Berg, K. Yamaguchi, T. Berg, K. Stratos, and H. Daumé III · 2012
Cited alongside, same era.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Cited alongside, same era.
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Closest in time.
Improving image-sentence embeddings using large weakly annotated photo collections
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik · 2014
Closest in time.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Closest in time.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and L. Fei-Fei · 2014
Closest in time.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Closest in time.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Closest in time.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Closest in time.