Fetching the paper…
Reading the bibliography…
Image and sentence matching has made great progress recently, but it remains challenging due to the large visual-semantic discrepancy.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Fisher kernels on visual vocabularies for image categorization
F. Perronnin and C. Dance · 2007
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Earlier work this paper cites.
Deep convolutional ranking for multilabel image annotation
Y. Gong, Y. Jia, T. Leung, A. Toshev, and S. Ioffe · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Cnn: Single-label to multi-label
Y. Wei, W. Xia, J. Huang, B. Ni, J. Dong, Y. Zhao, and S. Yan · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Earlier work this paper cites.
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. Lawrence Zitnick · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, et al · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and F.-F. Li · 2015
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2015
Cited alongside, same era.
Associating neural word embeddings with deep image representations using fisher vectors
B. Klein, G. Lev, G. Sadeh, and L. Wolf · 2015
Cited alongside, same era.
Multimodal convolutional neural networks for matching image and sentence
L. Ma, Z. Lu, L. Shang, and H. Li · 2015
Cited alongside, same era.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2015
Leveraging visual question answering for image-caption ranking
X. Lin and D. Parikh · 2016
Later among the works it cites.
Jointly modeling embedding and translation to bridge video and language
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui · 2016
Later among the works it cites.
Order-embeddings of images and language
I. Vendrov, R. Kiros, S. Fidler, and R. Urtasun · 2016
Later among the works it cites.
Cnn-rnn: A unified framework for multi-label image classification
J. Wang, Y. Yang, J. Mao, Z. Huang, C. Huang, and W. Xu · 2016
Later among the works it cites.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Later among the works it cites.
What value do explicit high level concepts have in vision to language problems?
Q. Wu, C. Shen, L. Liu, A. Dick, and A. van den Hengel · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. Plummer, L. Wang, C. Cervantes, J. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
Deep multiple instance learning for image classification and auto-annotation
J. Wu, Y. Yu, C. Huang, and K. Yu · 2015
Cited alongside, same era.
Deep correlation for matching images and text
F. Yan and K. Mikolajczyk · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Rnn fisher vectors for action recognition and image annotation
G. Lev, G. Sadeh, B. Klein, and L. Wolf · 2016
Cited alongside, same era.
Linking image and text with 2-way nets
A. Eisenschtat and L. Wolf · 2017
Closest in time.
Vse++: Improved visual-semantic embeddings
F. Faghri, D. J. Fleet, J. R. Kiros, and S. Fidler · 2017
Closest in time.
Instance-aware image and sentence matching with selective multimodal lstm
Y. Huang, W. Wang, and L. Wang · 2017
Closest in time.
Learning a recurrent residual fusion network for multimodal matching
Y. Liu, Y. Guo, E. M. Bakker, and M. S. Lew · 2017
Closest in time.
Dual attention networks for multimodal reasoning and matching
H. Nam, J.-W. Ha, and J. Kim · 2017
Closest in time.
Show and tell: Lessons learned from the 2015 mscoco image captioning challenge
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2017
Closest in time.