Fetching the paper…
Reading the bibliography…
Effective image and sentence matching depends on how to well measure their global visual-semantic similarity.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K. K. Paliwal · 1997
Earlier work this paper cites.
Top-down attentional guidance based on implicit learning of visual covariation
M. M. Chun and Y. Jiang · 1999
Earlier work this paper cites.
Generating typed dependency parses from phrase structure parses
M.-C. De Marneffe, B. MacCartney, C. D. Manning, et al · 2006
Earlier work this paper cites.
The role of context in object recognition
A. Oliva and A. Torralba · 2007
Earlier work this paper cites.
Fisher kernels on visual vocabularies for image categorization
F. Perronnin and C. Dance · 2007
Earlier work this paper cites.
Quantifying center bias of observers in free viewing of dynamic natural scenes
P.-H. Tseng, R. Carmi, I. G. Cameron, D. P. Munoz, and L. Itti · 2009
Earlier work this paper cites.
Scene and screen center bias early eye movements in scene viewing
M. Bindemann · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Cited alongside, same era.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and F.-F. Li · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Cited alongside, same era.
Multiple object recognition with visual attention
J. Ba, V. Mnih, and K. Kavukcuoglu · 2015
Associating neural word embeddings with deep image representations using fisher vectors
B. Klein, G. Lev, G. Sadeh, and L. Wolf · 2015
Later among the works it cites.
Multimodal convolutional neural networks for matching image and sentence
L. Ma, Z. Lu, L. Shang, and H. Li · 2015
Later among the works it cites.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2015
Later among the works it cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. Plummer, L. Wang, C. Cervantes, J. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Later among the works it cites.
Deep correlation for matching images and text
F. Yan and K. Mikolajczyk · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. Lawrence Zitnick · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, et al · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and F.-F. Li · 2015
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2015
Cited alongside, same era.
Skip-thought vectors
R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler · 2015
Cited alongside, same era.
Later among the works it cites.
Contextual lstm (clstm) models for large scale nlp tasks
S. Ghosh, O. Vinyals, B. Strope, S. Roy, T. Dean, and L. Heck · 2016
Closest in time.
Rnn fisher vectors for action recognition and image annotation
G. Lev, G. Sadeh, B. Klein, and L. Wolf · 2016
Closest in time.
Order-embeddings of images and language
I. Vendrov, R. Kiros, S. Fidler, and R. Urtasun · 2016
Closest in time.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2016
Closest in time.