Fetching the paper…
Reading the bibliography…
Image-language matching tasks have recently attracted a lot of attention in the computer vision field.
H. Hotelling, “Relations between two sets of variables,” Biometrika , vol. 28, p. 312–377, 1936
1936
Earlier work this paper cites.
J. Bromley, J. W. Bentz, L. Bottou, I. Guyon, Y. LeCun, C. Moore, E. Säckinger, and R. Shah, “Signature verification using a “siamese” time delay neural network,” International Journal of Pattern Recognition and Artificial Intelligence , vol. 7, no. 04, pp. 669–688, 1993
1993
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
D. R. Hardoon, S. Szedmak, and J. Shawe-Taylor, “Canonical correlation analysis: An overview with application to learning methods,” Neural computation , vol. 16, no. 12, pp. 2639–2664, 2004
2004
Earlier work this paper cites.
K. Q. Weinberger, J. Blitzer, and L. K. Saul, “Distance metric learning for large margin nearest neighbor classification,” in NIPS , 2005
2005
Earlier work this paper cites.
S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in CVPR , 2005
2005
Earlier work this paper cites.
B. Shaw and T. Jebara, “Structure preserving embedding,” in ICML , 2009
2009
Earlier work this paper cites.
T. Joachims, T. Finley, and C.-N. J. Yu, “Cutting-plane training of structural svms,” Machine Learning , vol. 77, no. 1, pp. 27–59, 2009
2009
Earlier work this paper cites.
F. Perronnin, J. Sanchez, and T. Mensink, “Improving the Fisher kernel for large-scale image classification,” in ECCV , 2010
2010
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in ICML , 2011
2011
Earlier work this paper cites.
J. Weston, S. Bengio, and N. Usunier, “Wsabie: Scaling up to large vocabulary image annotation,” in IJCAI , 2011
2011
Earlier work this paper cites.
B. Shaw, B. Huang, and T. Jebara, “Learning a distance metric from a network,” in NIPS , 2011
2011
Earlier work this paper cites.
M. Everingham, L. Van Gool, C. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge 2012,” 2011
2011
Earlier work this paper cites.
N. Srivastava and R. R. Salakhutdinov, “Multimodal learning with deep boltzmann machines,” in NIPS , 2012
2012
Earlier work this paper cites.
T. Mensink, J. Verbeek, F. Perronnin, and G. Csurka, “Metric learning for large scale image classification: Generalizing to new classes at near-zero cost,” in ECCV , 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in NIPS , 2012
2012
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in NIPS , 2013
2013
Earlier work this paper cites.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,” Journal of Artificial Intelligence Research , 2013
2013
Earlier work this paper cites.
G. Andrew, R. Arora, J. Bilmes, and K. Livescu, “Deep canonical correlation analysis,” in ICML , 2013
2013
Earlier work this paper cites.
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov et al. , “Devise: A deep visual-semantic embedding model,” in NIPS , 2013
2013
Earlier work this paper cites.
A. Karpathy, A. Joulin, and F. F. F. Li, “Deep fragment embeddings for bidirectional image sentence mapping,” in NIPS , 2014
2014
Earlier work this paper cites.
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik, “Improving image-sentence embeddings using large weakly annotated photo collections,” in ECCV , 2014
2014
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg, “Referitgame: Referring to objects in photographs of natural scenes.” in EMNLP , 2014
2014
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” Transactions of the Association for Computational Linguistics , vol. 2, pp. 67–78, 2014
2014
Cited alongside, same era.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV , 2014
2014
Cited alongside, same era.
Y. Gong, Q. Ke, M. Isard, and S. Lazebnik, “A multi-view embedding space for modeling internet images, tags, and their semantics,” IJCV , 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Z. Ma, Y. Lu, and D. Foster, “Finding linear structure in large datasets with scalable canonical correlation analysis,” ICML , 2015
2015
Later among the works it cites.
J. Mao, W. Xu, Y. Yang, J. Wang, and A. Yuille, “Deep captioning with multimodal recurrent neural networks (m-rnn),” ICLR , 2015
2015
Later among the works it cites.
X. Han, T. Leung, Y. Jia, R. Sukthankar, and A. C. Berg, “Matchnet: Unifying feature and metric learning for patch-based matching,” in CVPR , 2015
2015
Later among the works it cites.
E. Hoffer and N. Ailon, “Deep metric learning using triplet network,” ICLR , 2015
2015
Later among the works it cites.
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” CVPR , 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Kiros, R. Salakhutdinov, and R. S. Zemel, “Multimodal neural language models.” in ICML , 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2014
Cited alongside, same era.
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng, “Grounded compositional semantics for finding and describing images with sentences,” Transactions of the Association for Computational Linguistics , vol. 2, pp. 207–218, 2014
2014
Cited alongside, same era.
J. Hu, J. Lu, and Y.-P. Tan, “Discriminative deep metric learning for face verification in the wild,” in CVPR , 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
J. Wang, Y. Song, T. Leung, C. Rosenberg, J. Wang, J. Philbin, B. Chen, and Y. Wu, “Learning fine-grained image similarity with deep ranking,” in CVPR , 2014
2014
Cited alongside, same era.
C. L. Zitnick and P. Dollár, “Edge boxes: Locating object proposals from edges,” in ECCV , 2014
2014
Cited alongside, same era.
J. Ba, K. Swersky, S. Fidler, and R. Salakhutdinov, “Predicting deep zero-shot convolutional neural networks using textual descriptions,” ICCV , 2015
2015
Later among the works it cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” ICML , 2015
2015
Later among the works it cites.
R. Girshick, “Fast r-cnn,” in ICCV , 2015
2015
Later among the works it cites.
L. Ma, Z. Lu, L. Shang, and H. Li, “Multimodal convolutional neural networks for matching image and sentence,” ICCV , 2015
2015
Later among the works it cites.
J. Johnson, A. Karpathy, and L. Fei-Fei, “Densecap: Fully convolutional localization networks for dense captioning,” CVPR , 2016
2016
Later among the works it cites.
A. Jabri, A. Joulin, and L. van der Maaten, “Revisiting visual question answering baselines,” in ECCV , 2016
2016
Later among the works it cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in ECCV , 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
L. Wang, Y. Li, and S. Lazebnik, “Learning deep structure-preserving image-text embeddings,” CVPR , 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele, “Grounding of textual phrases in images by reconstruction,” ECCV , 2016
2016
Later among the works it cites.
B. Plummer, L. Wang, C. Cervantes, J. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” IJCV , 2016
2016
Later among the works it cites.
M. Wang, M. Azab, N. Kojima, R. Mihalcea, and J. Deng, “Structured matching for phrase localization,” in ECCV , 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
I. Vendrov, R. Kiros, S. Fidler, and R. Urtasun, “Order-embeddings of images and language,” ICLR , 2016
2016
Later among the works it cites.
A. Eisenschtat and L. Wolf, “Linking image and text with 2-way nets,” CVPR , 2017
2017
Closest in time.
L. Yu, H. Tan, M. Bansal, and T. L. Berg, “A joint speaker-listener-reinforcer model for referring expressions,” CVPR , 2017
2017
Closest in time.