Fetching the paper…
Reading the bibliography…
In recent years, cross-modal retrieval has drawn much attention due to the rapid growth of multimodal data.
Y. Chen, L. Wang, W. Wang, and Z. Zhang, “Continuum regression for cross-modal multimedia retrieval,” in International Conference on Image Processing . IEEE, 2012, pp. 1949–1952
1952
Earlier work this paper cites.
J. B. Tenenbaum and W. T. Freeman, “Separating style and content with bilinear models,” Neural Computation , vol. 12, no. 6, pp. 1247–1283, 2000
2000
Earlier work this paper cites.
D. Li, N. Dimitrova, M. Li, and I. K. Sethi, “Multimedia content processing through cross-modal association,” in International Conference on Multimedia . ACM, 2003, pp. 604–611
2003
Earlier work this paper cites.
D. M. Blei and M. I. Jordan, “Modeling annotated data,” in Conference on Research and Development in Informaion Retrieval . ACM, 2003, pp. 127–134
2003
Earlier work this paper cites.
D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent dirichlet allocation,” Journal of Machine Learning Research , vol. 3, pp. 993–1022, 2003
2003
Earlier work this paper cites.
D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International Journal of Computer Vision , vol. 60, no. 2, pp. 91–110, 2004
2004
Earlier work this paper cites.
R. Rosipal and N. Krämer, “Overview and recent advances in partial least squares,” in Subspace, latent structure and feature selection . Springer, 2006, pp. 34–51
2006
Earlier work this paper cites.
D. Lin and X. Tang, “Inter-modality face recognition,” in European Conference on Computer Vision . Springer, 2006, pp. 13–26
2006
Earlier work this paper cites.
D. Grangier and S. Bengio, “A discriminative kernel-based approach to rank images from text queries,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 30, no. 8, pp. 1371–1384, 2008
2008
Earlier work this paper cites.
L. Sun, S. Ji, and J. Ye, “A least squares formulation for canonical correlation analysis,” in International Conference on Machine learning . ACM, 2008, pp. 1024–1031
2008
Earlier work this paper cites.
Y. Weiss, A. Torralba, and R. Fergus, “Spectral hashing,” in Advances in Neural Information Processing Systems , 2009, pp. 1753–1760
2009
Earlier work this paper cites.
T.-S. Chua, J. Tang, R. Hong, H. Li, Z. Luo, and Y. Zheng, “NUS-WIDE: A real-world web image database from national university of singapore,” in International Conference on Image and Video Retrieval . ACM, 2009, p. 48
2009
Earlier work this paper cites.
J. Liu, C. Xu, and H. Lu, “Cross-media retrieval: state-of-the-art and open issues,” International Journal of Multimedia Intelligence and Security , vol. 1, no. 1, pp. 33–52, 2010
2010
Earlier work this paper cites.
N. Rasiwasia, J. Costa Pereira, E. Coviello, G. Doyle, G. R. Lanckriet, R. Levy, and N. Vasconcelos, “A new approach to cross-modal multimedia retrieval,” in International conference on Multimedia . ACM, 2010, pp. 251–260
2010
Earlier work this paper cites.
D. Putthividhy, H. T. Attias, and S. S. Nagarajan, “Topic regression multi-modal latent dirichlet allocation for image annotation,” in Computer Vision and Pattern Recognition . IEEE, 2010, pp. 3408–3415
2010
Earlier work this paper cites.
B. Bai, J. Weston, D. Grangier, R. Collobert, K. Sadamasa, Y. Qi, O. Chapelle, and K. Weinberger, “Learning to rank with (a lot of) word features,” Information Retrieval , vol. 13, no. 3, pp. 291–314, 2010
2010
Earlier work this paper cites.
R. Udupa and M. Khapra, “Improving the multilingual user experience of wikipedia using cross-language name search,” in Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics . Association for Computational Linguistics, 2010, pp. 492–500
2010
Earlier work this paper cites.
W. Wu, J. Xu, and H. Li, “Learning similarity function between objects in heterogeneous spaces,” Microsoft Research Technique Report , 2010
2010
Earlier work this paper cites.
D. Zhang, J. Wang, D. Cai, and J. Lu, “Self-taught hashing for fast similarity search,” in Conference on Research and Development in Information Retrieval . ACM, 2010, pp. 18–25
2010
Earlier work this paper cites.
J. Krapac, M. Allan, J. Verbeek, and F. Jurie, “Improving web image search results using query-relative classifiers,” in Computer Vision and Pattern Recognition , 2010, pp. 1094–1101
2010
Earlier work this paper cites.
V. Mahadevan, C. W. Wong, J. C. Pereira, T. Liu, N. Vasconcelos, and L. K. Saul, “Maximum covariance unfolding: Manifold learning for bimodal data,” in Advances in Neural Information Processing Systems , 2011, pp. 918–926
2011
Earlier work this paper cites.
Y. Jia, M. Salzmann, and T. Darrell, “Learning cross-modality similarity for multinomial data,” in International Conference on Computer Vision . IEEE, 2011, pp. 2407–2414
2011
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in International Conference on Machine Learning , 2011, pp. 689–696
2011
Earlier work this paper cites.
N. Quadrianto and C. H. Lampert, “Learning multi-view neighborhood preserving projections,” in International Conference on Machine Learning , 2011, pp. 425–432
2011
Earlier work this paper cites.
J. Weston, S. Bengio, and N. Usunier, “Wsabie: Scaling up to large vocabulary image annotation,” in International Joint Conference on Artificial Intelligence , vol. 11, 2011, pp. 2764–2770
2011
Earlier work this paper cites.
R. He, W.-S. Zheng, and B.-G. Hu, “Maximum correntropy criterion for robust face recognition,” Transactions on Pattern Analysis and Machine Intelligence , vol. 33, no. 8, pp. 1561–1576, 2011
2011
Earlier work this paper cites.
A. Li, S. Shan, X. Chen, and W. Gao, “Face recognition based on non-corresponding region matching,” in International Conference on Computer Vision . IEEE, 2011, pp. 1060–1067
2011
Earlier work this paper cites.
A. Sharma and D. W. Jacobs, “Bypassing synthesis: Pls for face recognition with pose, low-resolution and sketch,” in Computer Vision and Pattern Recognition . IEEE, 2011, pp. 593–600
2011
Earlier work this paper cites.
Y. Gong and S. Lazebnik, “Iterative quantization: A procrustean approach to learning binary codes,” in Computer Vision and Pattern Recognition . IEEE, 2011, pp. 817–824
2011
Earlier work this paper cites.
D. Zhang, F. Wang, and L. Si, “Composite hashing with multiple information sources,” in Conference on Research and Development in Information Retrieval . ACM, 2011, pp. 225–234
2011
Earlier work this paper cites.
J. Song, Y. Yang, Z. Huang, H. T. Shen, and R. Hong, “Multiple feature hashing for real-time large scale near-duplicate video retrieval,” in International Conference on Multimedia . ACM, 2011, pp. 423–432
2011
Earlier work this paper cites.
A. Sharma, A. Kumar, H. Daume, and D. W. Jacobs, “Generalized multiview analysis: A discriminative latent space,” in Computer Vision and Pattern Recognition . IEEE, 2012, pp. 2160–2167
2012
Earlier work this paper cites.
X. Shi and P. Yu, “Dimensionality reduction on heterogeneous feature space,” in International Conference on Data Mining , 2012, pp. 635–644
2012
Earlier work this paper cites.
N. Srivastava and R. R. Salakhutdinov, “Multimodal learning with deep boltzmann machines,” in Advances in Neural Information Processing Systems , 2012, pp. 2222–2230
2012
Earlier work this paper cites.
D. Zhai, H. Chang, S. Shan, X. Chen, and W. Gao, “Multiview metric learning with global consistency and local smoothness,” ACM Transactions on Intelligent Systems and Technology , vol. 3, no. 3, 2012
2012
Earlier work this paper cites.
Y. Zhen and D.-Y. Yeung, “Co-regularized hashing for multimodal data,” in Advances in Neural Information Processing Systems , 2012, pp. 1376–1384
2012
Earlier work this paper cites.
Y. Zhen and D.-Y. Yeung, “A probabilistic model for multimodal hash function learning,” in International Conference on Knowledge Discovery and Data Mining . ACM, 2012, pp. 940–948
2012
Earlier work this paper cites.
A. Mignon and F. Jurie, “CMML: a new metric learning approach for cross modal matching,” in Asian Conference on Computer Vision , 2012, pp. 14–pages
2012
Cited alongside, same era.
S. J. Hwang and K. Grauman, “Reading between the lines: Object localization using implicit cues from image tags,” Transactions on Pattern Analysis and Machine Intelligence , vol. 34, no. 6, pp. 1145–1158, 2012
2012
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Neural Information Processing Systems , 2012, pp. 1106–1114
2012
Cited alongside, same era.
C. Xu, D. Tao, and C. Xu, “A survey on multi-view learning,” arXiv preprint arXiv:1304.5634 , 2013
2013
Cited alongside, same era.
D. Zhang and W.-J. Li, “Large-scale supervised multimodal hashing with semantic correlation maximization,” in AAAI Conference on Artificial Intelligence , 2014, pp. 2177–2183
2014
Later among the works it cites.
Y. Gong, Q. Ke, M. Isard, and S. Lazebnik, “A multi-view embedding space for modeling internet images, tags, and their semantics,” International Journal of Computer Vision , vol. 106, no. 2, pp. 210–233, 2014
2014
Later among the works it cites.
J. Costa Pereira, E. Coviello, G. Doyle, N. Rasiwasia, G. R. Lanckriet, R. Levy, and N. Vasconcelos, “On the role of correlation and abstraction in cross-modal multimedia retrieval,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 36, no. 3, pp. 521–535, 2014
2014
Later among the works it cites.
Y. Peter, L. Alice, H. Micah, and H. Julia, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” in Transactions of the Association for Computational Linguistics , 2014, pp. 67–78
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
X. Zhai, Y. Peng, and J. Xiao, “Heterogeneous metric learning with joint graph regularization for cross-media retrieval,” in AAAI Conference on Artificial Intelligence , 2013, pp. 1198–1204
2013
Cited alongside, same era.
Z. Yuan, J. Sang, Y. Liu, and C. Xu, “Latent feature learning in social media network,” in International Conference on Multimedia . ACM, 2013, pp. 253–262
2013
Cited alongside, same era.
X. Lu, F. Wu, S. Tang, Z. Zhang, X. He, and Y. Zhuang, “A low rank structural large margin method for cross-modal ranking,” in Conference on Research and Development in Information Retrieval . ACM, 2013, pp. 433–442
2013
Cited alongside, same era.
F. Wu, X. Lu, Z. Zhang, S. Yan, Y. Rui, and Y. Zhuang, “Cross-media semantic representation via bi-directional learning to rank,” in International Conference on Multimedia . ACM, 2013, pp. 877–886
2013
Cited alongside, same era.
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov et al. , “Devise: A deep visual-semantic embedding model,” in Advances in Neural Information Processing Systems , 2013, pp. 2121–2129
2013
Cited alongside, same era.
X. Mao, B. Lin, D. Cai, X. He, and J. Pei, “Parallel field alignment for cross media retrieval,” in International Conference on Multimedia . ACM, 2013, pp. 897–906
2013
Cited alongside, same era.
Y. T. Zhuang, Y. F. Wang, F. Wu, Y. Zhang, and W. M. Lu, “Supervised coupled dictionary learning with group structures for multi-modal retrieval,” in AAAI Conference on Artificial Intelligence , 2013
2013
Cited alongside, same era.
2014
Later among the works it cites.
2014
Later among the works it cites.
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and F. Li, “Large-scale video classification with convolutional neural networks,” in Computer Vision and Pattern Recognition , 2014, pp. 1725–1732
2014
Later among the works it cites.
2014
Later among the works it cites.
A. Karpathy and F. Li, “Deep visual-semantic alignments for generating image descriptions,” in Computer Vision and Pattern Recognition , 2015, pp. 3128–3137
2015
Later among the works it cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Computer Vision and Pattern Recognition , 2015, pp. 3156–3164
2015
Later among the works it cites.
X. Chen and C. L. Zitnick, “Mind’s eye: A recurrent visual representation for image caption generation,” in Computer Vision and Pattern Recognition , 2015, pp. 2422–2431
2015
Later among the works it cites.
X. Jia, E. Gavves, B. Fernando, and T. Tuytelaars, “Guiding the long-short term memory model for image caption generation,” in International Conference on Computer Vision , 2015, pp. 2407–2415
2015
Later among the works it cites.
Y. Ushiku, M. Yamaguchi, Y. Mukuta, and T. Harada, “Common subspace for model and similarity: Phrase learning for caption generation from images,” in International Conference on Computer Vision , 2015, pp. 2668–2676
2015
Later among the works it cites.
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille, “Deep captioning with multimodal recurrent neural networks (m-rnn),” 2015
2015
Later among the works it cites.
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko, “Translating videos to natural language using deep recurrent neural networks,” North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pp. 1494–1504, 2015
2015
Later among the works it cites.
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, T. Darrell, and K. Saenko, “Long-term recurrent convolutional networks for visual recognition and description,” in Computer Vision and Pattern Recognition , 2015, pp. 2625–2634
2015
Later among the works it cites.
F. Yan and K. Mikolajczyk, “Deep correlation for matching images and text,” 2015, pp. 3441–3450
2015
Later among the works it cites.
R. Xu, C. Xiong, W. Chen, and C. J. J., “Jointly modeling deep video and compositional text to bridge vision and language in a unified framework,” in AAAI Conference on Artificial Intelligence , 2015, pp. 2346–2352
2015
Later among the works it cites.
J. Wang, Y. He, C. Kang, S. Xiang, and C. Pan, “Image-text cross-modal retrieval via modality-specific feature learning,” in International Conference on Multimedia Retrieval , 2015, pp. 347–354
2015
Later among the works it cites.
T. Yao, T. Mei, and C.-W. Ngo, “Learning query and image similarities with ranking canonical correlation analysis,” in International Conference on Computer Vision , 2015, pp. 28–36
2015
Later among the works it cites.
X. Jiang, F. Wu, X. Li, Z. Zhao, W. Lu, S. Tang, and Y. Zhuang, “Deep compositional cross-modal learning to rank via local-global alignment,” in International Conference on Multimedia . ACM, 2015, pp. 69–78
2015
Later among the works it cites.
C. Wang, H. Yang, and C. Meinel, “Deep semantic mapping for cross-modal retrieval,” in International Conference on Tools with Artificial Intelligence , 2015, pp. 234–241
2015
Later among the works it cites.
D. Wang, P. Cui, M. Ou, and W. Zhu, “Learning compact hash codes for multimodal representations using orthogonal deep structure,” IEEE Transactions on Multimedia , vol. 17, no. 9, pp. 1404–1416, 2015
2015
Later among the works it cites.
B. Wu, Q. Yang, W. Zheng, Y. Wang, and J. Wang, “Quantized correlation hashing for fast cross-modal search,” in International Joint Conference on Artificial Intelligence , 2015, pp. 3946–3952
2015
Later among the works it cites.
Z. Lin, G. Ding, M. Hu, and J. Wang, “Semantics-preserving hashing for cross-view retrieval,” in Computer Vision and Pattern Recognition , 2015, pp. 3864–3872
2015
Later among the works it cites.
Y. Hua, H. Tian, A. Cai, and P. Shi, “Cross-modal correlation learning with deep convolutional architecture,” in Visual Communications and Image Processing , 2015, pp. 1–4
2015
Later among the works it cites.
V. Ranjan, N. Rasiwasia, and C. V. Jawahar, “Multi-label cross-modal retrieval,” 2015, pp. 4094–4102
2015
Later among the works it cites.
Z. Li, W. Lu, E. Bao, and W. Xing, “Learning a semantic space by deep network for cross-media retrieval,” in International Conference on Distributed Multimedia Systems , 2015, pp. 199–203
2015
Later among the works it cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale visual recognition,” in International Conference on Learning Representations , 2015
2015
Later among the works it cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Computer Vision and Pattern Recognition , 2015, pp. 1–9
2015
Later among the works it cites.
W. Wang, X. Yang, B. C. Ooi, D. Zhang, and Y. Zhuang, “Effective deep learning-based multi-modal retrieval,” International Journal on Very Large Data Bases , vol. 25, no. 1, pp. 79–101, 2016
2016
Closest in time.
2016
Closest in time.
Y. Cao, M. Long, J. Wang, and H. Zhu, “Correlation autoencoder hashing for supervised cross-modal search,” in International Conference on Multimedia Retrieval , 2016
2016
Closest in time.
2016
Closest in time.
L. Castrejon, Y. Aytar, C. Vondrick, H. Pirsiavash, and A. Torralba, “Learning aligned cross-modal representations from weakly aligned data,” in Computer Vision and Pattern Recognition , 2016
2016
Closest in time.
G. Ding, Y. Guo, and J. Zhou, “Collective matrix factorization hashing for multimodal data,” in Computer Vision and Pattern Recognition . IEEE, 2014, pp. 2083–2090
2090
Closest in time.
K. Wang, R. He, W. Wang, L. Wang, and T. Tan, “Learning coupled feature spaces for cross-modal matching,” in International Conference on Computer Vision . IEEE, 2013, pp. 2088–2095
2095
Closest in time.