Fetching the paper…
Reading the bibliography…
It is known that the inconsistent distribution and representation of different modalities, such as image and text, cause the heterogeneity gap that makes it challenging to correlate such heterogeneous data.
H. Hotelling, “Relations between two sets of variates,” Biometrika , pp. 321–377, 1936
1936
Earlier work this paper cites.
H. McGurk and J. MacDonald, “Hearing lips and seeing voices,” Nature , vol. 264, no. 5588, pp. 746–748, 1976
1976
Earlier work this paper cites.
D. Li, N. Dimitrova, M. Li, and I. K. Sethi, “Multimedia content processing through cross-modal association,” in ACM International Conference on Multimedia (ACM-MM) , 2003, pp. 604–611
2003
Earlier work this paper cites.
D. R. Hardoon, S. Szedmák, and J. Shawe-Taylor, “Canonical correlation analysis: An overview with application to learning methods,” Neural Computation , vol. 16, no. 12, pp. 2639–2664, 2004
2004
Earlier work this paper cites.
Y. Yang, Y. Zhuang, F. Wu, and Y. Pan, “Harmonizing hierarchical manifolds for multimedia document semantics understanding and cross-media retrieval,” IEEE Transactions on Multimedia (TMM) , vol. 10, no. 3, pp. 437–446, 2008
2008
Earlier work this paper cites.
Y. Zhuang, Y. Yang, and F. Wu, “Mining semantic correlation of heterogeneous multimedia data for cross-media retrieval,” IEEE Transactions on Multimedia (TMM) , vol. 10, no. 2, pp. 221–229, 2008
2008
Earlier work this paper cites.
N. Rasiwasia, J. Costa Pereira, E. Coviello, G. Doyle, G. R. Lanckriet, R. Levy, and N. Vasconcelos, “A new approach to cross-modal multimedia retrieval,” in ACM International Conference on Multimedia (ACM-MM) , 2010, pp. 251–260
2010
Earlier work this paper cites.
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier, “Collecting image annotations using amazon’s mechanical turk,” in NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk , 2010, pp. 139–147
2010
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in International Conference on Machine Learning (ICML) , 2011, pp. 689–696
2011
Earlier work this paper cites.
H. G. Krizhevsky A, Sutskever I, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (NIPS) , 2012
2012
Earlier work this paper cites.
J. Kim, J. Nam, and I. Gurevych, “Learning semantics with deep belief network for cross-language information retrieval,” in International Committee on Computational Linguistic (ICCL) , 2012, pp. 579–588
2012
Earlier work this paper cites.
N. Srivastava and R. Salakhutdinov, “Learning representations for multimodal data with deep belief nets,” in International Conference on Machine Learning (ICML) Workshop , 2012
2012
Earlier work this paper cites.
X. Zhai, Y. Peng, and J. Xiao, “Heterogeneous metric learning with joint graph regularization for cross-media retrieval,” in AAAI Conference on Artificial Intelligence (AAAI) , 2013, pp. 1198–1204
2013
Earlier work this paper cites.
G. Andrew, R. Arora, J. A. Bilmes, and K. Livescu, “Deep canonical correlation analysis,” in International Conference on Machine Learning (ICML) , 2013, pp. 1247–1255
2013
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in Advances in Neural Information Processing Systems (NIPS) , 2013, pp. 3111–3119
2013
Earlier work this paper cites.
Y. Gong, Q. Ke, M. Isard, and S. Lazebnik, “A multi-view embedding space for modeling internet images, tags, and their semantics,” International Journal of Computer Vision (IJCV) , vol. 106, no. 2, pp. 210–233, 2014
2014
Cited alongside, same era.
F. Feng, X. Wang, and R. Li, “Cross-modal retrieval with correspondence autoencoder,” in ACM International Conference on Multimedia (ACM-MM) , 2014, pp. 7–16
2014
Cited alongside, same era.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems (NIPS) , 2014, pp. 2672–2680
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Y. Peng, X. Huang, and J. Qi, “Cross-media shared representation by hierarchical learning with multiple deep networks,” in International Joint Conference on Artificial Intelligence (IJCAI) , 2016, pp. 3846–3853
2016
Later among the works it cites.
C. Finn, I. Goodfellow, and S. Levine, “Unsupervised learning for physical interaction through video prediction,” in Advances in Neural Information Processing Systems (NIPS) , 2016, pp. 64–72
2016
Later among the works it cites.
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee, “Generative adversarial text to image synthesis,” in International Conference on Machine Learning (ICML) , 2016, pp. 1060–1069
2016
Later among the works it cites.
K. Wang, R. He, L. Wang, W. Wang, and T. Tan, “Joint feature selection and subspace learning for cross-modal retrieval,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , vol. 38, no. 10, pp. 2010–2023, 2016
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Zhai, Y. Peng, and J. Xiao, “Learning cross-media joint representation with sparse and semi-supervised regularization,” IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) , vol. 24, pp. 965–978, 2014
2014
Cited alongside, same era.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in International Conference on Learning Representations (ICLR) , 2014
2014
Cited alongside, same era.
Y. Kim, “Convolutional neural networks for sentence classification,” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2014, pp. 1746–1751
2014
Cited alongside, same era.
L. Pang, S. Zhu, and C. Ngo, “Deep multimodal learning for affective analysis and retrieval,” IEEE Transactions on Multimedia (TMM) , vol. 17, no. 11, pp. 2008–2020, 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
V. Ranjan, N. Rasiwasia, and C. V. Jawahar, “Multi-label cross-modal retrieval,” in IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 4094–4102
2015
Cited alongside, same era.
S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster R-CNN: towards real-time object detection with region proposal networks,” in Advances in Neural Information Processing Systems (NIPS) , 2015, pp. 91–99
2015
Cited alongside, same era.
D. Wang, P. Cui, M. Ou, and W. Zhu, “Deep multimodal hashing with orthogonal regularization,” in International Joint Conference on Artificial Intelligence (IJCAI) , 2015, pp. 2291–2297
2015
Cited alongside, same era.
2016
Later among the works it cites.
X. Wang and A. Gupta, “Generative image modeling using style and structure adversarial networks,” in European Conference on Computer Vision (ECCV) , 2016, pp. 318–335
2016
Later among the works it cites.
2016
Later among the works it cites.
S. E. Reed, Z. Akata, S. Mohan, S. Tenka, B. Schiele, and H. Lee, “Learning what and where to draw,” in Advances in Neural Information Processing Systems (NIPS) , 2016, pp. 217–225
2016
Later among the works it cites.
Y. Peng, W. Zhu, Y. Zhao, C. Xu, Q. Huang, H. Lu, Q. Zheng, T. Huang, and W. Gao, “Cross-media analysis and reasoning: advances and directions,” Frontiers of Information Technology & Electronic Engineering , vol. 18, no. 1, pp. 44–57, 2017
2017
Closest in time.
Y. Peng, X. Huang, and Y. Zhao, “An overview of cross-media retrieval: Concepts, methodologies, benchmarks and challenges,” IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) , 2017
2017
Closest in time.
L. Zhang, B. Ma, G. Li, Q. Huang, and Q. Tian, “Cross-modal retrieval using multi-ordered discriminative structured subspace learning,” IEEE Transactions on Multimedia (TMM) , vol. 19, no. 6, pp. 1220–1233, 2017
2017
Closest in time.
2017
Closest in time.
Y. Peng, J. Qi, X. Huang, and Y. Yuan, “Ccl: Cross-modal correlation learning with multi-grained fusion by hierarchical network,” IEEE Transactions on Multimedia (TMM) , 2017
2017
Closest in time.
H. Zhang, T. Xu, H. Li, S. Zhang, X. Huang, X. Wang, and D. Metaxas, “Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks,” in IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 1–8
2017
Closest in time.
Y. Wei, Y. Zhao, C. Lu, S. Wei, L. Liu, Z. Zhu, and S. Yan, “Cross-modal retrieval with CNN visual features: A new baseline,” IEEE Transactions on Cybernetics (TCYB) , vol. 47, no. 2, pp. 449–460, 2017
2017
Closest in time.