Fetching the paper…
Reading the bibliography…
With the exponential surge in diverse multi-modal data, traditional uni-modal retrieval methods struggle to meet the needs of users seeking access to data across various modalities.
Z. Wu, S. Song, A. Khosla, F. Yu, L. Zhang, X. Tang, and J. Xiao, “3d shapenets: A deep representation for volumetric shapes,” in IEEE CVPR , 2015, pp. 1912–1920
1920
Earlier work this paper cites.
Y. Chen, S. Wang, J. Lu, Z. Chen, Z. Zhang, and Z. Huang, “Local graph convolutional networks for cross-modal hashing,” in ACM MM , 2021, pp. 1921–1928
1928
Earlier work this paper cites.
L. Bensabath, M. Petrovich, and G. Varol, “A cross-dataset study for text-based 3d human motion retrieval,” in IEEE CVPR , 2024, pp. 1932–1940
1940
Earlier work this paper cites.
Y. Chen, L. Wang, W. Wang, and Z. Zhang, “Continuum regression for cross-modal multimedia retrieval,” in IEEE ICIP , 2012, pp. 1949–1952
1952
Earlier work this paper cites.
Y. Song and M. Soleymani, “Polysemous visual-semantic embedding for cross-modal retrieval,” in IEEE CVPR , 2019, pp. 1979–1988
1988
Earlier work this paper cites.
A. Smeulders, M. Worring, S. Santini, A. Gupta, and R. Jain, “Content-based image retrieval at the end of the early years,” IEEE TPAMI , vol. 22, no. 12, pp. 1349–1380, 2000
2000
Earlier work this paper cites.
D. M. Blei and M. I. Jordan, “Modeling annotated data,” in ACM SIGIR , 2003, pp. 127–134
2003
Earlier work this paper cites.
G. Griffin, A. Holub, and P. Perona, “Caltech-256 object category dataset,” 2007
2007
Earlier work this paper cites.
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE TNN , vol. 20, no. 1, pp. 61–80, 2008
2008
Earlier work this paper cites.
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol, “Extracting and composing robust features with denoising autoencoders,” in ICML , 2008, pp. 1096–1103
2008
Earlier work this paper cites.
G. Chechik, E. Ie, M. Rehn, S. Bengio, and D. Lyon, “Large-scale content-based audio retrieval from text queries,” in ACM ICMIR , 2008, pp. 105–112
2008
Earlier work this paper cites.
M. J. Huiskes and M. S. Lew, “The mir flickr retrieval evaluation,” in ACM ICMIR , 2008, pp. 39–43
2008
Earlier work this paper cites.
B. C. Russell, A. Torralba, K. P. Murphy, and W. T. Freeman, “Labelme: a database and web-based tool for image annotation,” IJCV , vol. 77, pp. 157–173, 2008
2008
Earlier work this paper cites.
T.-S. Chua, J. Tang, R. Hong, H. Li, Z. Luo, and Y. Zheng, “Nus-wide: a real-world web image database from national university of singapore,” in ACM ICIVR , 2009, pp. 1–9
2009
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE TKDE , vol. 22, no. 10, pp. 1345–1359, 2009
2009
Earlier work this paper cites.
N. Rasiwasia, J. C. Pereira, E. Coviello, G. Doyle, G. R. G. Lanckriet, R. Levy, and N. Vasconcelos, “A new approach to cross-modal multimedia retrieval,” in ACM MM , 2010, pp. 251–260
2010
Earlier work this paper cites.
D. Putthividhya, H. T. Attias, and S. S. Nagarajan, “Topic regression multi-modal latent dirichlet allocation for image annotation,” in IEEE CVPR , 2010, pp. 3408–3415
2010
Earlier work this paper cites.
H. Jégou, M. Douze, C. Schmid, and P. Pérez, “Aggregating local descriptors into a compact image representation,” in IEEE CVPR , 2010, pp. 3304–3311
2010
Earlier work this paper cites.
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier, “Collecting image annotations using amazon’s mechanical turk,” in NAACL-HLT , 2010, pp. 139–147
2010
Earlier work this paper cites.
J. Krapac, M. Allan, J. Verbeek, and F. Juried, “Improving web image search results using query-relative classifiers,” in IEEE CVPR , 2010, pp. 1094–1101
2010
Earlier work this paper cites.
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier, “Collecting image annotations using amazon’s mechanical turk,” in NAACL-HLT , 2010, pp. 139–147
2010
Earlier work this paper cites.
H. J. Escalante, C. A. Hernández, J. A. Gonzalez, A. López-López, M. Montes, E. F. Morales, L. E. Sucar, L. Villasenor, and M. Grubinger, “The segmented and annotated iapr tc-12 benchmark,” CVIU , vol. 114, no. 4, pp. 419–428, 2010
2010
Earlier work this paper cites.
L. Zhang and Y. Zhang, “Interactive retrieval based on faceted feedback,” in ACM SIGIR , 2010, pp. 363–370
2010
Earlier work this paper cites.
S. Kumar and R. Udupa, “Learning hash functions for cross-view similarity search,” in IJCAI , 2011, pp. 1360–1365
2011
Earlier work this paper cites.
Y. Jia, M. Salzmann, and T. Darrell, “Learning cross-modality similarity for multinomial data,” in IEEE ICCV , 2011, pp. 2407–2414
2011
Earlier work this paper cites.
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie, “The caltech-ucsd birds-200-2011 dataset,” 2011
2011
Earlier work this paper cites.
A. Sharma, A. Kumar, H. D. III, and D. W. Jacobs, “Generalized multiview analysis: A discriminative latent space,” in IEEE CVPR , 2012, pp. 2160–2167
2012
Earlier work this paper cites.
Y. Zhen and D. Yeung, “Co-regularized hashing for multimodal data,” in NIPS , 2012, pp. 1385–1393
2012
Earlier work this paper cites.
Y. Li, J. Chen, and L. Feng, “Dealing with uncertainty: A survey of theories and practices,” IEEE TKDE , vol. 25, no. 11, pp. 2463–2482, 2012
2012
Earlier work this paper cites.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,” JAIR , vol. 47, pp. 853–899, 2013
2013
Earlier work this paper cites.
G. Andrew, R. Arora, J. A. Bilmes, and K. Livescu, “Deep canonical correlation analysis,” in ICML , vol. 28, 2013, pp. 1247–1255
2013
Earlier work this paper cites.
Y. Zhuang, Y. Wang, F. Wu, Y. Zhang, and W. Lu, “Supervised coupled dictionary learning with group structures for multi-modal retrieval,” in AAAI , vol. 27, no. 1, 2013, pp. 1070–1076
2013
Earlier work this paper cites.
X. Mao, B. Lin, D. Cai, X. He, and J. Pei, “Parallel field alignment for cross media retrieval,” in ACM MM , 2013, pp. 897–906
2013
Earlier work this paper cites.
X. Zhu, Z. Huang, H. T. Shen, and X. Zhao, “Linear cross-modal hashing for efficient multimedia search,” in ACM MM , 2013, pp. 143–152
2013
Earlier work this paper cites.
J. Song, Y. Yang, Y. Yang, Z. Huang, and H. T. Shen, “Inter-media hashing for large-scale retrieval from heterogeneous data sources,” in ACM SIGMOD , 2013, pp. 785–796
2013
Earlier work this paper cites.
M. Rastegari, J. Choi, S. Fakhraei, H. D. III, and L. S. Davis, “Predictable dual-view hashing,” in ICML , vol. 28, 2013, pp. 1328–1336
2013
Earlier work this paper cites.
M. Ou, P. Cui, F. Wang, J. Wang, W. Zhu, and S. Yang, “Comparing apples to oranges: a scalable solution with heterogeneous hashing,” in ACM KDD , 2013, pp. 230–238
2013
Earlier work this paper cites.
V. K. Vavilapalli, A. C. Murthy, C. Douglas, S. Agarwal, M. Konar, R. Evans, T. Graves, J. Lowe, H. Shah, S. Seth et al. , “Apache hadoop yarn: Yet another resource negotiator,” in ACM SOCC , 2013, pp. 1–16
2013
Earlier work this paper cites.
F. Feng, X. Wang, and R. Li, “Cross-modal retrieval with correspondence autoencoder,” in ACM MM , 2014, pp. 7–16
2014
Earlier work this paper cites.
J. Masci, M. M. Bronstein, A. M. Bronstein, and J. Schmidhuber, “Multimodal similarity-preserving hashing,” IEEE TPAMI , vol. 36, no. 4, pp. 824–830, 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in NIPS , vol. 27, 2014
2014
Earlier work this paper cites.
Y. Gong, Q. Ke, M. Isard, and S. Lazebnik, “A multi-view embedding space for modeling internet images, tags, and their semantics,” IJCV , vol. 106, no. 2, pp. 210–233, 2014
2014
Earlier work this paper cites.
X. Zhai, Y. Peng, and J. Xiao, “Learning cross-media joint representation with sparse and semisupervised regularization,” IEEE TCSVT , vol. 24, no. 6, pp. 965–978, 2014
2014
Earlier work this paper cites.
Y. Wang, F. Wu, J. Song, X. Li, and Y. Zhuang, “Multi-modal mutual topic reinforce modeling for cross-media retrieval,” in ACM MM , 2014, pp. 307–316
2014
Earlier work this paper cites.
R. Liao, J. Zhu, and Z. Qin, “Nonparametric bayesian upstream supervised multi-modal topic models,” in ACM WSDM , 2014, pp. 493–502
2014
Earlier work this paper cites.
J. Zhou, G. Ding, and Y. Guo, “Latent semantic sparse hashing for cross-modal similarity search,” in ACM SIGIR , 2014, pp. 415–424
2014
Earlier work this paper cites.
F. Wu, Z. Yu, Y. Yang, S. Tang, Y. Zhang, and Y. Zhuang, “Sparse multi-modal hashing,” IEEE TMM , vol. 16, no. 2, pp. 427–439, 2014
2014
Earlier work this paper cites.
Y. Hu, Z. Jin, H. Ren, D. Cai, and X. He, “Iterative multi-view hashing for cross media indexing,” in ACM MM , 2014, pp. 527–536
2014
Earlier work this paper cites.
Z. Yu, F. Wu, Y. Yang, Q. Tian, J. Luo, and Y. Zhuang, “Discriminative coupled dictionary hashing for fast cross-media retrieval,” in ACM SIGIR , 2014, pp. 395–404
2014
Earlier work this paper cites.
D. Zhang and W. Li, “Large-scale supervised multimodal hashing with semantic correlation maximization,” in AAAI , 2014, pp. 2177–2183
2014
Earlier work this paper cites.
Y. Zhuang, Z. Yu, W. Wang, F. Wu, S. Tang, and J. Shao, “Cross-media hashing with neural networks,” in ACM MM , 2014, pp. 901–904
2014
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” TACL , vol. 2, pp. 67–78, 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV , 2014, pp. 740–755
2014
Earlier work this paper cites.
M. I. Jordan and T. M. Mitchell, “Machine learning: Trends, perspectives, and prospects,” Science , vol. 349, no. 6245, pp. 255–260, 2015
2015
Earlier work this paper cites.
T. Yao, T. Mei, and C. Ngo, “Learning query and image similarities with ranking canonical correlation analysis,” in IEEE ICCV , 2015, pp. 28–36
2015
Earlier work this paper cites.
J. Wang, Y. He, C. Kang, S. Xiang, and C. Pan, “Image-text cross-modal retrieval via modality-specific feature learning,” in ACM ICMR , 2015, pp. 347–354
2015
Earlier work this paper cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in IEEE CVPR , 2015, pp. 3128–3137
2015
Earlier work this paper cites.
V. Ranjan, N. Rasiwasia, and C. V. Jawahar, “Multi-label cross-modal retrieval,” in IEEE ICCV , 2015, pp. 4094–4102
2015
Earlier work this paper cites.
D. Wang, X. Gao, X. Wang, and L. He, “Semantic topic multimodal hashing for cross-media retrieval,” in IJCAI , 2015, pp. 3890–3896
2015
Earlier work this paper cites.
G. Irie, H. Arai, and Y. Taniguchi, “Alternating co-quantization for cross-modal hashing,” in IEEE ICCV , 2015, pp. 1886–1894
2015
Earlier work this paper cites.
D. Wang, P. Cui, M. Ou, and W. Zhu, “Learning compact hash codes for multimodal representations using orthogonal deep structure,” IEEE TMM , vol. 17, no. 9, pp. 1404–1416, 2015
2015
Earlier work this paper cites.
B. Wu, Q. Yang, W. Zheng, Y. Wang, and J. Wang, “Quantized correlation hashing for fast cross-modal search,” in IJCAI , 2015, pp. 3946–3952
2015
Earlier work this paper cites.
N. Kriegeskorte, “Deep neural networks: a new framework for modeling biological vision and brain information processing,” Annual Review of Vision Science , vol. 1, pp. 417–446, 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y. Peng, X. Huang, and J. Qi, “Cross-media shared representation by hierarchical learning with multiple deep networks,” in IJCAI , 2016, pp. 3846–3853
2016
Earlier work this paper cites.
W. Wang, X. Yang, B. C. Ooi, D. Zhang, and Y. Zhuang, “Effective deep learning-based multi-modal retrieval,” The VLDB Journal , vol. 25, no. 1, pp. 79–101, 2016
2016
Earlier work this paper cites.
C. Deng, X. Tang, J. Yan, W. Liu, and X. Gao, “Discriminative dictionary learning with common label alignment for cross-modal retrieval,” IEEE TMM , vol. 18, no. 2, pp. 208–218, 2016
2016
Earlier work this paper cites.
K. Wang, R. He, L. Wang, W. Wang, and T. Tan, “Joint feature selection and subspace learning for cross-modal retrieval,” IEEE TPAMI , vol. 38, no. 10, pp. 2010–2023, 2016
2016
Earlier work this paper cites.
Y. Wei, Y. Zhao, Z. Zhu, S. Wei, Y. Xiao, J. Feng, and S. Yan, “Modality-dependent cross-media retrieval,” ACM TIST , vol. 7, no. 4, pp. 57:1–57:13, 2016
2016
Earlier work this paper cites.
L. Zhang, B. Ma, G. Li, Q. Huang, and Q. Tian, “Pl-ranking: A novel ranking method for cross-modal retrieval,” in ACM MM , 2016, pp. 1355–1364
2016
Earlier work this paper cites.
M. Long, Y. Cao, J. Wang, and P. S. Yu, “Composite correlation quantization for efficient multimodal retrieval,” in ACM SIGIR , 2016, pp. 579–588
2016
Earlier work this paper cites.
J. Tang, K. Wang, and L. Shao, “Supervised matrix factorization hashing for cross-modal retrieval,” IEEE TIP , vol. 25, no. 7, pp. 3157–3166, 2016
2016
Earlier work this paper cites.
X. Xu, “Dictionary learning based hashing for cross-modal retrieval,” in ACM MM , 2016, pp. 177–181
2016
Earlier work this paper cites.
D. Wang, X. Gao, X. Wang, L. He, and B. Yuan, “Multimodal discriminative binary embedding for large-scale cross-modal retrieval,” IEEE TIP , vol. 25, no. 10, pp. 4540–4554, 2016
2016
Earlier work this paper cites.
T. Yan, X. Xu, S. Guo, Z. Huang, and X. Wang, “Supervised robust discrete multimodal hashing for cross-media retrieval,” in ACM CIKM , 2016, pp. 1271–1280
2016
Earlier work this paper cites.
Y. Cao, M. Long, J. Wang, and H. Zhu, “Correlation autoencoder hashing for supervised cross-modal search,” in ACM ICMR , 2016, pp. 197–204
2016
Earlier work this paper cites.
Y. Cao, M. Long, J. Wang, Q. Yang, and P. S. Yu, “Deep visual-semantic hashing for cross-modal retrieval,” in ACM KDD , 2016, pp. 1445–1454
2016
Earlier work this paper cites.
L. Xie, J. Shen, and L. Zhu, “Online cross-modal hashing for web image retrieval,” in AAAI , 2016, pp. 294–300
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
D.-D. Le, S. Phan, V.-T. Nguyen, B. Renoust, T. A. Nguyen, V.-N. Hoang, T. D. Ngo, M.-T. Tran, Y. Watanabe, M. Klinkigt et al. , “Nii-hitachi-uit at trecvid 2016.” in TRECVID , vol. 25, 2016
2016
Earlier work this paper cites.
M. Foteini, M. Anastasia, G. Damianos, M. Theodoros, K. Vagia, I. Anastasia, and S. Symeonidis, “Iti-certh participation in trecvid 2016,” in TRECVID 2016 Workshop , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Xu, T. Mei, T. Yao, and Y. Rui, “Msr-vtt: A large video description dataset for bridging video and language,” in IEEE CVPR , 2016, pp. 5288–5296
2016
Earlier work this paper cites.
B. Qu, X. Li, D. Tao, and X. Lu, “Deep semantic understanding of high resolution remote sensing image,” in IEEE CITS , 2016, pp. 1–5
2016
Earlier work this paper cites.
X. Xu, A. Dehghani, D. Corrigan, S. Caulfield, and D. Moloney, “Convolutional neural network for 3d object recognition using volumetric representation,” in IEEE SPLINE , 2016, pp. 1–5
2016
Earlier work this paper cites.
K. Weiss, T. M. Khoshgoftaar, and D. Wang, “A survey of transfer learning,” Journal of Big Data , vol. 3, no. 1, pp. 1–40, 2016
2016
Earlier work this paper cites.
M. Zaharia, R. S. Xin, P. Wendell, T. Das, M. Armbrust, A. Dave, X. Meng, J. Rosen, S. Venkataraman, M. J. Franklin et al. , “Apache spark: a unified engine for big data processing,” Communications of the ACM , vol. 59, no. 11, pp. 56–65, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Wang, T. Zhang, N. Sebe, H. T. Shen et al. , “A survey on learning to hash,” IEEE TPAMI , vol. 40, no. 4, pp. 769–790, 2017
2017
Earlier work this paper cites.
Y. Peng, X. Huang, and Y. Zhao, “An overview of cross-media retrieval: Concepts, methodologies, benchmarks, and challenges,” IEEE TCSVT , vol. 28, no. 9, pp. 2372–2385, 2017
2017
Earlier work this paper cites.
Y. Peng, J. Qi, X. Huang, and Y. Yuan, “Ccl: Cross-modal correlation learning with multigrained fusion by hierarchical network,” IEEE TMM , vol. 20, no. 2, pp. 405–420, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Y. Liu, Y. Guo, E. M. Bakker, and M. S. Lew, “Learning a recurrent residual fusion network for multimodal matching,” in IEEE ICCV , 2017, pp. 4107–4116
2017
Earlier work this paper cites.
Y. Huang, W. Wang, and L. Wang, “Instance-aware image and sentence matching with selective multimodal LSTM,” in IEEE CVPR , 2017, pp. 7254–7262
2017
Earlier work this paper cites.
A. Eisenschtat and L. Wolf, “Linking image and text with 2-way nets,” in IEEE CVPR , 2017, pp. 1855–1865
2017
Earlier work this paper cites.
J. Wu, Z. Lin, and H. Zha, “Joint latent subspace learning and regression for cross-modal retrieval,” in ACM SIGIR , 2017, pp. 917–920
2017
Earlier work this paper cites.
Y. Wu, S. Wang, and Q. Huang, “Online asymmetric similarity learning for cross-modal retrieval,” in IEEE CVPR , 2017, pp. 3984–3993
2017
Earlier work this paper cites.
B. Wang, Y. Yang, X. Xu, A. Hanjalic, and H. T. Shen, “Adversarial cross-modal retrieval,” in ACM MM , 2017, pp. 154–162
2017
Earlier work this paper cites.
X. Li, D. Hu, and F. Nie, “Deep binary reconstruction for cross-modal hashing,” in ACM MM , 2017, pp. 1398–1406
2017
Earlier work this paper cites.
X. Xu, F. Shen, Y. Yang, H. T. Shen, and X. Li, “Learning discriminative binary codes for large-scale cross-modal retrieval,” IEEE TIP , vol. 26, no. 5, pp. 2494–2507, 2017
2017
Earlier work this paper cites.
K. Li, G. Qi, J. Ye, and K. A. Hua, “Linear subspace ranking hashing for cross-modal retrieval,” IEEE TPAMI , vol. 39, no. 9, pp. 1825–1838, 2017
2017
Earlier work this paper cites.
E. Yang, C. Deng, W. Liu, X. Liu, D. Tao, and X. Gao, “Pairwise relationship guided deep hashing for cross-modal retrieval,” in AAAI , 2017, pp. 1618–1625
2017
Earlier work this paper cites.
Q. Jiang and W. Li, “Deep cross-modal hashing,” in IEEE CVPR , 2017, pp. 3270–3278
2017
Earlier work this paper cites.
Y. Cao, M. Long, J. Wang, and S. Liu, “Collective deep quantization for efficient cross-modal retrieval,” in AAAI , 2017, pp. 3974–3980
2017
Earlier work this paper cites.
D. Mandal, K. N. Chaudhury, and S. Biswas, “Generalized semantic preserving hashing for n-label cross-modal retrieval,” in IEEE CVPR , 2017, pp. 2633–2641
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. Ueki, K. Hirakawa, K. Kikuchi, T. Ogawa, and T. Kobayashi, “Waseda_meisei at trecvid 2017: Ad-hoc video search.” in TRECVID , 2017
2017
Earlier work this paper cites.
F. Markatopoulou, D. Galanopoulos, V. Mezaris, and I. Patras, “Query and keyframe representations for ad-hoc video search,” in ACM ICMR , 2017, pp. 407–411
2017
Earlier work this paper cites.
Y. Yu, H. Ko, J. Choi, and G. Kim, “End-to-end concept word detection for video captioning, retrieval, and question answering,” in IEEE CVPR , 2017, pp. 3165–3173
2017
Earlier work this paper cites.
A. Araujo and B. Girod, “Large-scale video retrieval using image queries,” IEEE TCSVT , vol. 28, no. 6, pp. 1406–1420, 2017
2017
Earlier work this paper cites.
Y. Peng, X. Huang, and Y. Zhao, “An overview of cross-media retrieval: Concepts, methodologies, benchmarks, and challenges,” IEEE TCSVT , vol. 28, no. 9, pp. 2372–2385, 2017
2017
Earlier work this paper cites.
A. Salvador, N. Hynes, Y. Aytar, J. Marin, F. Ofli, I. Weber, and A. Torralba, “Learning cross-modal embeddings for cooking recipes and food images,” in IEEE CVPR , 2017, pp. 3020–3028
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” IJCV , vol. 123, pp. 32–73, 2017
2017
Earlier work this paper cites.
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, “Dense-captioning events in videos,” in IEEE ICCV , 2017, pp. 706–715
2017
Earlier work this paper cites.
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele, “Movie description,” IJCV , vol. 123, pp. 94–120, 2017
2017
Earlier work this paper cites.
X. Lu, B. Wang, X. Zheng, and X. Li, “Exploring models and data for remote sensing image caption generation,” IEEE TGRS , vol. 56, no. 4, pp. 2183–2195, 2017
2017
Earlier work this paper cites.
T. Kim, M. Cha, H. Kim, J. K. Lee, and J. Kim, “Learning to discover cross-domain relations with generative adversarial networks,” in ICML , 2017, pp. 1857–1865
2017
Earlier work this paper cites.
H. B. McMahan, “A survey of algorithms and analysis for adaptive online learning,” JMLR , vol. 18, no. 1, pp. 3117–3166, 2017
2017
Earlier work this paper cites.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE TPAMI , vol. 41, no. 2, pp. 423–443, 2018
2018
Earlier work this paper cites.
S. Pouyanfar, S. Sadiq, Y. Yan, H. Tian, Y. Tao, M. P. Reyes, M.-L. Shyu, S.-C. Chen, and S. S. Iyengar, “A survey on deep learning: Algorithms, techniques, and applications,” ACM Computing Surveys , vol. 51, no. 5, pp. 1–36, 2018
2018
Earlier work this paper cites.
N. C. Mithun, R. Panda, E. E. Papalexakis, and A. K. Roy-Chowdhury, “Webly supervised joint embedding for cross-modal image-text retrieval,” in ACM MM , 2018, pp. 1856–1864
2018
Earlier work this paper cites.
Y. Zhan, J. Yu, Z. Yu, R. Zhang, D. Tao, and Q. Tian, “Comprehensive distance-preserving autoencoders for cross-modal retrieval,” in ACM MM , 2018, pp. 1137–1145
2018
Earlier work this paper cites.
J. Wehrmann and R. C. Barros, “Bidirectional retrieval made simple,” in IEEE CVPR , 2018, pp. 7718–7726
2018
Earlier work this paper cites.
J. Qi, Y. Peng, and Y. Yuan, “Cross-media multi-level alignment with relation attention network,” in IJCAI , 2018, pp. 892–898
2018
Earlier work this paper cites.
M. Engilberge, L. Chevallier, P. Pérez, and M. Cord, “Finding beans in burgers: Deep semantic-visual embedding with localization,” in IEEE CVPR , 2018, pp. 3984–3993
2018
Earlier work this paper cites.
J. Gu, J. Cai, S. R. Joty, L. Niu, and G. Wang, “Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models,” in IEEE CVPR , 2018, pp. 7181–7189
2018
Earlier work this paper cites.
Y. Huang, Q. Wu, C. Song, and L. Wang, “Learning semantic concepts and order for image and sentence matching,” in IEEE CVPR , 2018, pp. 6163–6171
2018
Earlier work this paper cites.
Y. Peng, J. Qi, and Y. Yuan, “Modality-specific cross-modal similarity measurement with recurrent attention network,” IEEE TIP , vol. 27, no. 11, pp. 5585–5599, 2018
2018
Earlier work this paper cites.
Y. Wu, S. Wang, and Q. Huang, “Learning semantic structure-preserved embeddings for cross-modal retrieval,” in ACM MM , 2018, pp. 825–833
2018
Earlier work this paper cites.
D. Wang, Q. Wang, and X. Gao, “Robust and flexible discrete hashing for cross-modal similarity search,” IEEE TCSVT , vol. 28, no. 10, pp. 2703–2715, 2018
2018
Earlier work this paper cites.
F. Zheng, Y. Tang, and L. Shao, “Hetero-manifold regularisation for cross-modal hashing,” IEEE TPAMI , vol. 40, no. 5, pp. 1059–1071, 2018
2018
Earlier work this paper cites.
G. Wu, Z. Lin, J. Han, L. Liu, G. Ding, B. Zhang, and J. Shen, “Unsupervised deep hashing via binary latent factor models for large-scale cross-modal retrieval,” in IJCAI , 2018, pp. 2854–2860
2018
Earlier work this paper cites.
J. Zhang, Y. Peng, and M. Yuan, “Unsupervised generative adversarial cross-modal hashing,” in AAAI , 2018, pp. 539–546
2018
Earlier work this paper cites.
X. Liu, X. Nie, W. Zeng, C. Cui, L. Zhu, and Y. Yin, “Fast discrete cross-modal hashing with regressing from semantic labels,” in ACM MM , 2018, pp. 1662–1669
2018
Earlier work this paper cites.
C. Deng, Z. Chen, X. Liu, X. Gao, and D. Tao, “Triplet-based deep hashing network for cross-modal retrieval,” IEEE TIP , vol. 27, no. 8, pp. 3893–3903, 2018
2018
Cited alongside, same era.
C. Li, C. Deng, N. Li, W. Liu, X. Gao, and D. Tao, “Self-supervised adversarial hashing networks for cross-modal retrieval,” in IEEE CVPR , 2018, pp. 4242–4251
2018
Cited alongside, same era.
X. Xu, J. Song, H. Lu, Y. Yang, F. Shen, and Z. Huang, “Modal-adversarial semantic learning network for extendable cross-modal retrieval,” in ACM ICMR , 2018, pp. 46–54
2018
Cited alongside, same era.
N. C. Mithun, J. Li, F. Metze, and A. K. Roy-Chowdhury, “Learning joint embedding with multimodal cues for cross-modal video-text retrieval,” in ACM ICMR , 2018, pp. 19–27
2018
Cited alongside, same era.
X. Hao, W. Zhang, D. Wu, F. Zhu, and B. Li, “Listen and look: Multi-modal aggregation and co-attention network for video-audio retrieval,” in IEEE ICME , 2022, pp. 1–6
2022
Later among the works it cites.
A. S. Koepke, A.-M. Oncescu, J. F. Henriques, Z. Akata, and S. Albanie, “Audio retrieval with natural language queries: A benchmark study,” IEEE TMM , vol. 25, pp. 2675–2685, 2022
2022
Later among the works it cites.
Z. Wang, Z. Gao, X. Xu, Y. Luo, Y. Yang, and H. T. Shen, “Point to rectangle matching for image text retrieval,” in ACM MM , 2022, pp. 4977–4986
2022
Later among the works it cites.
M. Cheng, Y. Sun, L. Wang, X. Zhu, K. Yao, J. Chen, G. Song, J. Han, J. Liu, E. Ding, and J. Wang, “Vista: Vision and scene text aggregation for cross-modal retrieval,” in IEEE CVPR , 2022, pp. 5174–5183
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
J. Dong, X. Li, and C. G. Snoek, “Predicting visual features from text for image and video caption retrieval,” IEEE TMM , vol. 20, no. 12, pp. 3377–3388, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in ACL , 2018, pp. 2556–2565
2018
Cited alongside, same era.
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu, “Audio-visual event localization in unconstrained videos,” in ECCV , 2018, pp. 247–263
2018
Cited alongside, same era.
Y. Zhou, Z. Wang, C. Fang, T. Bui, and T. L. Berg, “Visual to sound: Generating natural sound for videos in the wild,” in IEEE CVPR , 2018, pp. 3550–3558
2018
Cited alongside, same era.
X. Sun, J. Wu, X. Zhang, Z. Zhang, C. Zhang, T. Xue, J. B. Tenenbaum, and W. T. Freeman, “Pix3d: Dataset and methods for single-image 3d shape modeling,” in IEEE CVPR , 2018, pp. 2974–2983
2018
Cited alongside, same era.
K. Lee, X. Chen, G. Hua, H. Hu, and X. He, “Stacked cross attention for image-text matching,” in ECCV , 2018, pp. 212–228
2018
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Blattmann, R. Rombach, K. Oktay, J. Müller, and B. Ommer, “Retrieval-augmented diffusion models,” in NIPS , vol. 35, 2022, pp. 15 309–15 324
2022
Later among the works it cites.
Q. Cui, B. Zhou, Y. Guo, W. Yin, H. Wu, O. Yoshie, and Y. Chen, “Contrastive vision-language pre-training with limited resources,” in ECCV , 2022, pp. 236–253
2022
Later among the works it cites.
2022
Later among the works it cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in IEEE CVPR , 2022, pp. 16 000–16 009
2022
Later among the works it cites.
M. Dang, W. H. Xiang, L. Y. Fen, and T. N. Nguyen, “Explainable artificial intelligence: a comprehensive review,” Artificial Intelligence Review , vol. 55, no. 5, pp. 3503–3568, 2022
2022
Later among the works it cites.
Z. Liu, F. Chen, J. Xu, W. Pei, and G. Lu, “Image-text retrieval with cross-modal semantic importance consistency,” IEEE TCSVT , vol. 33, no. 5, pp. 2465–2476, 2023
2023
Closest in time.
F.-L. Chen, D.-Z. Zhang, M.-L. Han, X.-Y. Chen, J. Shi, S. Xu, and B. Xu, “Vlp: A survey on vision-language pre-training,” Machine Intelligence Research , vol. 20, no. 1, pp. 38–56, 2023
2023
Closest in time.
Y. Wang, Y. Su, W. Li, J. Xiao, X. Li, and A.-A. Liu, “Dual-path rare content enhancement network for image and text matching,” IEEE TCSVT , vol. 33, no. 10, pp. 6144–6158, 2023
2023
Closest in time.
C. Liu, Y. Zhang, H. Wang, W. Chen, F. Wang, Y. Huang, Y.-D. Shen, and L. Wang, “Efficient token-guided image-text retrieval with consistent multimodal contrastive training,” IEEE TIP , vol. 32, pp. 3622–3633, 2023
2023
Closest in time.
Z. Fu, Z. Mao, Y. Song, and Y. Zhang, “Learning semantic relationship among instances for image-text matching,” in IEEE CVPR , 2023, pp. 15 159–15 168
2023
Closest in time.
X. Wang, L. Li, Z. Li, X. Wang, X. Zhu, C. Wang, J. Huang, and Y. Xiao, “AGREE: aligning cross-modal entities for image-text retrieval upon vision-language pre-trained models,” in ACM WSDM , 2023, pp. 456–464
2023
Closest in time.
D. Jiang and M. Ye, “Cross-modal implicit relation reasoning and aligning for text-to-image person retrieval,” in IEEE CVPR , 2023, pp. 2787–2797
2023
Closest in time.
X. Zheng, Z. Wang, S. Li, K. Xu, T. Zhuang, Q. Liu, and X. Zeng, “Make: Vision-language pre-training based product retrieval in taobao search,” in ACM WWW , 2023, pp. 356–360
2023
Closest in time.
S. He, W. Wang, Z. Wang, X. Xu, Y. Yang, X. Wang, and H. T. Shen, “Category alignment adversarial learning for cross-modal retrieval,” IEEE TKDE , vol. 35, no. 5, pp. 4527–4538, 2023
2023
Closest in time.
Z. Li, H. Lu, H. Fu, Z. Wang, and G. Gu, “Adaptive adversarial learning based cross-modal retrieval,” EAAI , vol. 123, p. 106439, 2023
2023
Closest in time.
X. Tang, Y. Wang, J. Ma, X. Zhang, F. Liu, and L. Jiao, “Interacting-enhancing feature transformer for cross-modal remote-sensing image and text retrieval,” IEEE TGRS , vol. 61, pp. 1–15, 2023
2023
Closest in time.
R.-C. Tu, J. Jiang, Q. Lin, C. Cai, S. Tian, H. Wang, and W. Liu, “Unsupervised cross-modal hashing with modality-interaction,” IEEE TCSVT , vol. 33, no. 9, pp. 5296–5308, 2023
2023
Closest in time.
R.-C. Tu, X.-L. Mao, Q. Lin, W. Ji, W. Qin, W. Wei, and H. Huang, “Unsupervised cross-modal hashing via semantic text mining,” IEEE TMM , vol. 25, pp. 8946–8957, 2023
2023
Closest in time.
L. Zhu, X. Wu, J. Li, Z. Zhang, W. Guan, and H. T. Shen, “Work together: Correlation-identity reconstruction hashing for unsupervised cross-modal retrieval,” IEEE TKDE , vol. 35, no. 9, pp. 8838–8851, 2023
2023
Closest in time.
X. Xia, G. Dong, F. Li, L. Zhu, and X. Ying, “When clip meets cross-modal hashing retrieval: A new strong baseline,” Information Fusion , p. 101968, 2023
2023
Closest in time.
H. Li, C. Zhang, X. Jia, Y. Gao, and C. Chen, “Adaptive label correlation based asymmetric discrete hashing for cross-modal retrieval,” IEEE TKDE , vol. 35, no. 2, pp. 1185–1199, 2023
2023
Closest in time.
X. Liu, H. Zeng, Y. Shi, J. Zhu, C.-H. Hsia, and K.-K. Ma, “Deep cross-modal hashing based on semantic consistent ranking,” IEEE TMM , vol. 25, pp. 9530–9542, 2023
2023
Closest in time.
W. Tan, L. Zhu, J. Li, H. Zhang, and J. Han, “Teacher-student learning: Efficient hierarchical message aggregation hashing for cross-modal retrieval,” IEEE TMM , vol. 25, pp. 4520–4532, 2023
2023
Closest in time.
Y. Liu, Q. Wu, Z. Zhang, J. Zhang, and G. Lu, “Multi-granularity interactive transformer hashing for cross-modal retrieval,” in ACM MM , 2023, pp. 893–902
2023
Closest in time.
C. Zhang, H. Li, Y. Gao, and C. Chen, “Weakly-supervised enhanced semantic-aware hashing for cross-modal retrieval,” IEEE TKDE , vol. 35, no. 6, pp. 6475–6488, 2023
2023
Closest in time.
P. Hu, Z. Huang, D. Peng, X. Wang, and X. Peng, “Cross-modal retrieval with partially mismatched pairs,” IEEE TPAMI , vol. 45, no. 8, pp. 9595–9610, 2023
2023
Closest in time.
K. Luo, C. Zhang, H. Li, X. Jia, and C. Chen, “Adaptive marginalized semantic hashing for unpaired cross-modal retrieval,” IEEE TMM , vol. 25, pp. 9082–9095, 2023
2023
Closest in time.
H. Zhu, Y. Wei, X. Liang, C. Zhang, and Y. Zhao, “Ctp: Towards vision-language continual pretraining via compatible momentum contrast and topology preservation,” in IEEE ICCV , 2023, pp. 22 257–22 267
2023
Closest in time.
G. Song, X. Tan, and M. Yang, “Deep continual hashing with gradient-aware memory for cross-modal retrieval,” PR , vol. 137, p. 109276, 2023
2023
Closest in time.
Y. Xia, H. Huang, J. Zhu, and Z. Zhao, “Achieving cross modal generalization with multimodal unified representation,” in AAAI , vol. 36, 2023, pp. 63 529–63 541
2023
Closest in time.
D. Zhang, X. Wu, and G. Chen, “ONION: online semantic autoencoder hashing for cross-modal retrieval,” ACM TIST , vol. 14, no. 2, pp. 27:1–27:18, 2023
2023
Closest in time.
K. Jiang, W. K. Wong, X. Fang, J. Li, J. Qin, and S. Xie, “Random online hashing for cross-modal retrieval,” IEEE TNNLS , pp. 1–15, 2023
2023
Closest in time.
J. Li, F. Li, L. Zhu, H. Cui, and J. Li, “Prototype-guided knowledge transfer for federated unsupervised cross-modal hashing,” in ACM MM , 2023, pp. 1013–1022
2023
Closest in time.
X. Xu, J. Sun, Z. Cao, Y. Zhang, X. Zhu, and H. T. Shen, “Tfun: Trilinear fusion network for ternary image-text retrieval,” Information Fusion , vol. 91, pp. 327–337, 2023
2023
Closest in time.
A. Ray, F. Radenovic, A. Dubey, B. Plummer, R. Krishna, and K. Saenko, “Cola: A benchmark for compositional text-to-image retrieval,” in NIPS , vol. 36, 2023, pp. 46 433–46 445
2023
Closest in time.
D. Wang, C. Zhang, Q. Wang, Y. Tian, L. He, and L. Zhao, “Hierarchical semantic structure preserving hashing for cross-modal retrieval,” IEEE TMM , vol. 25, pp. 1217–1229, 2023
2023
Closest in time.
T. Wang, L. Zhu, Z. Zhang, H. Zhang, and J. Han, “Targeted adversarial attack against deep cross-modal hashing retrieval,” IEEE TCSVT , vol. 33, no. 10, pp. 6159–6172, 2023
2023
Closest in time.
L. Zhu, T. Wang, J. Li, Z. Zhang, J. Shen, and X. Wang, “Efficient query-based black-box attack against cross-modal hashing retrieval,” ACM TOIS , vol. 41, no. 3, pp. 1–25, 2023
2023
Closest in time.
P. Zhang, G. Bai, H. Yin, and Z. Huang, “Proactive privacy-preserving learning for cross-modal retrieval,” ACM TOIS , vol. 41, no. 2, pp. 35:1–35:23, 2023
2023
Closest in time.
J. Zhu, P. Zeng, L. Gao, G. Li, D. Liao, and J. Song, “Complementarity-aware space learning for video-text retrieval,” IEEE TCSVT , vol. 33, no. 8, pp. 4362–4374, 2023
2023
Closest in time.
S. Huang, B. Gong, Y. Pan, J. Jiang, Y. Lv, Y. Li, and D. Wang, “Vop: Text-video co-operative prompt tuning for cross-modal retrieval,” in IEEE CVPR , 2023, pp. 6565–6574
2023
Closest in time.
W. Wu, H. Luo, B. Fang, J. Wang, and W. Ouyang, “Cap4video: What can auxiliary captions do for text-video retrieval?” in IEEE CVPR , 2023, pp. 10 704–10 713
2023
Closest in time.
X. Shen, X. Zhang, X. Yang, Y. Zhan, L. Lan, J. Dong, and H. Wu, “Semantics-enriched cross-modal alignment for complex-query video moment retrieval,” in ACM MM , 2023, pp. 4109–4118
2023
Closest in time.
S. Zhao, L. Xu, Y. Liu, and S. Du, “Multi-grained representation learning for cross-modal retrieval,” in ACM SIGIR , 2023, pp. 2194–2198
2023
Closest in time.
Y. Xin, D. Yang, and Y. Zou, “Improving text-audio retrieval by text-aware attention pooling and prior matrix revised loss,” in IEEE ICASSP , 2023, pp. 1–5
2023
Closest in time.
T. Nakatsuka, M. Hamasaki, and M. Goto, “Content-based music-image retrieval using self- and cross-modal feature embedding memory,” in IEEE WACV , 2023, pp. 2174–2184
2023
Closest in time.
J. Huang, Y. Chen, S. Xiong, and X. Lu, “Self-supervision interactive alignment for remote sensing image-audio retrieval,” IEEE TGRS , vol. 61, pp. 1–14, 2023
2023
Closest in time.
Y. Chen, J. Huang, S. Xiong, and X. Lu, “Fine aligned discriminative hashing for remote sensing image-audio retrieval,” IEEE TGRS , vol. 61, pp. 1–12, 2023
2023
Closest in time.
Y. Kuang and X. Fan, “Collaborative audio-visual event localization based on sequential decision and cross-modal consistency,” in IEEE ICASSP , 2023, pp. 1–5
2023
Closest in time.
D. Zeng, J. Wu, G. Hattori, R. Xu, and Y. Yu, “Learning explicit and implicit dual common subspaces for audio-visual cross-modal retrieval,” ACM TOMM , vol. 19, no. 2s, pp. 1–23, 2023
2023
Closest in time.
J. Zhang, Y. Yu, S. Tang, J. Wu, and W. Li, “Variational autoencoder with cca for audio-visual cross-modal retrieval,” ACM TOMM , vol. 19, no. 3s, pp. 1–21, 2023
2023
Closest in time.
Y. Feng, H. Zhu, D. Peng, X. Peng, and P. Hu, “Rono: robust discriminative learning with noisy labels for 2d-3d cross-modal retrieval,” in IEEE CVPR , 2023, pp. 11 610–11 619
2023
Closest in time.
M. Petrovich, M. J. Black, and G. Varol, “Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis,” in IEEE ICCV , 2023, pp. 9488–9497
2023
Closest in time.
Y. Jiang, C. Hua, Y. Feng, and Y. Gao, “Hierarchical set-to-set representation for 3-d cross-modal retrieval,” IEEE TNNLS , pp. 1–13, 2023
2023
Closest in time.
C. Ge, J. Wang, Q. Qi, H. Sun, T. Xu, and J. Liao, “Scene-level sketch-based image retrieval with minimal pairwise supervision,” in AAAI , vol. 37, no. 1, 2023, pp. 650–657
2023
Closest in time.
Z. Wang, Z. Gao, K. Guo, Y. Yang, X. Wang, and H. T. Shen, “Multilateral semantic relations modeling for image text retrieval,” in IEEE CVPR , 2023, pp. 2830–2839
2023
Closest in time.
L. Wen, Y. Wang, D. Zhang, and G. Chen, “Visual matching is enough for scene text retrieval,” in ACM WSDM , 2023, pp. 447–455
2023
Closest in time.
J. Pan, Q. Ma, and C. Bai, “Reducing semantic confusion: Scene-aware aggregation network for remote sensing cross-modal retrieval,” in ACM ICMR , 2023, pp. 398–406
2023
Closest in time.
T. van Sonsbeek and M. Worring, “X-tra: Improving chest x-ray tasks with cross-modal retrieval augmentation,” in IPMI , 2023, pp. 471–482
2023
Closest in time.
F. Lin, M. Li, D. Li, T. Hospedales, Y.-Z. Song, and Y. Qi, “Zero-shot everything sketch-based image retrieval, and in explainable style,” in IEEE CVPR , 2023, pp. 23 349–23 358
2023
Closest in time.
A. Chaudhuri, A. K. Bhunia, Y.-Z. Song, and A. Dutta, “Data-free sketch-based image retrieval,” in IEEE CVPR , 2023, pp. 12 084–12 093
2023
Closest in time.
R. Ramos, B. Martins, D. Elliott, and Y. Kementchedjhieva, “Smallcap: lightweight image captioning prompted with retrieval augmentation,” in IEEE CVPR , 2023, pp. 2840–2849
2023
Closest in time.
W. Lin, J. Chen, J. Mei, A. Coca, and B. Byrne, “Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering,” in NIPS , vol. 36, 2023, pp. 22 820–22 840
2023
Closest in time.
Z. Gao, J. Wang, G. Yu, Z. Yan, C. Domeniconi, and J. Zhang, “Long-tail cross modal hashing,” in AAAI , vol. 37, no. 6, 2023, pp. 7642–7650
2023
Closest in time.
K. Zhou, Z. Liu, Y. Qiao, T. Xiang, and C. C. Loy, “Domain generalization: A survey,” IEEE TPAMI , vol. 45, no. 4, pp. 4396–4415, 2023
2023
Closest in time.
E. Fleisig, A. Amstutz, C. Atalla, S. L. Blodgett, H. Daumé III, A. Olteanu, E. Sheng, D. Vann, and H. Wallach, “Fairprism: evaluating fairness-related harms in text generation,” in ACL , 2023, pp. 6231–6251
2023
Closest in time.
2023
Closest in time.
J. Chen, H. Dong, X. Wang, F. Feng, M. Wang, and X. He, “Bias and debias in recommender system: A survey and future directions,” ACM TOIS , vol. 41, no. 3, pp. 1–39, 2023
2023
Closest in time.
W. X. Zhao, J. Liu, R. Ren, and J.-R. Wen, “Dense text retrieval based on pretrained language models: A survey,” ACM TOIS , vol. 42, no. 4, pp. 1–60, 2024
2024
Closest in time.
Z. Wang, Z. Gao, M. Han, Y. Yang, and H. T. Shen, “Estimating the semantics via sector embedding for image-text retrieval,” IEEE TMM , pp. 1–12, 2024
2024
Closest in time.
G. Xiong, M. Meng, T. Zhang, D. Zhang, and Y. Zhang, “Reference-aware adaptive network for image-text matching,” IEEE TCSVT , pp. 1–1, 2024
2024
Closest in time.
K. Pham, C. Huynh, S.-N. Lim, and A. Shrivastava, “Composing object relations and attributes for image-text matching,” in IEEE CVPR , 2024, pp. 14 354–14 363
2024
Closest in time.
L. Bao, L. Wei, W. Zhou, L. Liu, L. Xie, H. Li, and Q. Tian, “Multi-granularity matching transformer for text-based person search,” IEEE TMM , vol. 26, pp. 4281–4293, 2024
2024
Closest in time.
Y. Zhang, Z. Ji, D. Wang, Y. Pang, and X. Li, “User: Unified semantic enhancement with momentum contrast for image-text retrieval,” IEEE TIP , vol. 33, pp. 595–609, 2024
2024
Closest in time.
Z. Fu, L. Zhang, H. Xia, and Z. Mao, “Linguistic-aware patch slimming framework for fine-grained cross-modal alignment,” in IEEE CVPR , 2024, pp. 26 307–26 316
2024
Closest in time.
2024
Closest in time.
Z. Wang, X. Xu, J. Wei, N. Xie, Y. Yang, and H. T. Shen, “Semantics disentangling for cross-modal retrieval,” IEEE TIP , vol. 33, pp. 2226–2237, 2024
2024
Closest in time.
L. Zhang, L. Chen, C. Zhou, X. Li, F. Yang, and Z. Yi, “Weighted graph-structured semantics constraint network for cross-modal retrieval,” IEEE TMM , vol. 26, pp. 1551–1564, 2024
2024
Closest in time.
X. Liu, J. Li, X. Nie, X. Zhang, S. Wang, and Y. Yin, “Scalable unsupervised hashing via exploiting robust cross-modal consistency,” IEEE TBD , vol. 10, no. 4, pp. 514–527, 2024
2024
Closest in time.
H. Cai, B. Zhang, J. Li, B. Hu, and J. Chen, “Unsupervised dual hashing coding (udc) on semantic tagging and sample content for cross-modal retrieval,” IEEE TMM , vol. 26, pp. 9109–9120, 2024
2024
Closest in time.
M. Liang, J. Du, Z. Liang, Y. Xing, W. Huang, and Z. Xue, “Self-supervised multi-modal knowledge graph contrastive hashing for cross-modal search,” in AAAI , vol. 38, no. 12, 2024, pp. 13 744–13 753
2024
Closest in time.
J. Li, W. K. Wong, L. Jiang, X. Fang, S. Xie, and Y. Xu, “Ckdh: Clip-based knowledge distillation hashing for cross-modal retrieval,” IEEE TCSVT , vol. 34, no. 7, pp. 6530–6541, 2024
2024
Closest in time.
Z. Yang, X. Deng, L. Guo, and J. Long, “Asymmetric supervised fusion-oriented hashing for cross-modal retrieval,” IEEE TCYB , vol. 54, no. 2, pp. 851–864, 2024
2024
Closest in time.
F. Yang, M. Han, F. Ma, Y. Liu, X. Ding, and D. Tong, “Disperse asymmetric subspace relation hashing for cross-modal retrieval,” IEEE TCSVT , vol. 34, no. 1, pp. 603–617, 2024
2024
Closest in time.
Y. Sun, Z. Ren, P. Hu, D. Peng, and X. Wang, “Hierarchical consensus hashing for cross-modal retrieval,” IEEE TMM , vol. 26, pp. 824–836, 2024
2024
Closest in time.
Y. Wang, Y.-W. Zhan, Z.-D. Chen, X. Luo, and X.-S. Xu, “Multiple information embedded hashing for large-scale cross-modal retrieval,” IEEE TCSVT , vol. 34, no. 6, pp. 5118–5131, 2024
2024
Closest in time.
X. Liang, E. Yang, Y. Yang, and C. Deng, “Multi-relational deep hashing for cross-modal search,” IEEE TIP , vol. 33, pp. 3009–3020, 2024
2024
Closest in time.
Z. Hu, Y.-m. Cheung, M. Li, and W. Lan, “Cross-modal hashing method with properties of hamming space: A new perspective,” IEEE TPAMI , pp. 1–15, 2024
2024
Closest in time.
M. Meng, J. Sun, J. Liu, J. Yu, and J. Wu, “Semantic disentanglement adversarial hashing for cross-modal retrieval,” IEEE TCSVT , vol. 34, no. 3, pp. 1914–1926, 2024
2024
Closest in time.
Y. Bai, Z. Shu, J. Yu, Z. Yu, and X.-J. Wu, “Proxy-based graph convolutional hashing for cross-modal retrieval,” IEEE TBD , vol. 10, no. 4, pp. 371–385, 2024
2024
Closest in time.
C. Bai, C. Zeng, Q. Ma, and J. Zhang, “Graph convolutional network discrete hashing for cross-modal retrieval,” IEEE TNNLS , vol. 35, no. 4, pp. 4756–4767, 2024
2024
Closest in time.
Y. Huo, Q. Qin, J. Dai, L. Wang, W. Zhang, L. Huang, and C. Wang, “Deep semantic-aware proxy hashing for multi-label cross-modal retrieval,” IEEE TCSVT , vol. 34, no. 1, pp. 576–589, 2024
2024
Closest in time.
Y. Huo, Q. Qin, W. Zhang, L. Huang, and J. Nie, “Deep hierarchy-aware proxy hashing with self-paced learning for cross-modal retrieval,” IEEE TKDE , pp. 1–14, 2024
2024
Closest in time.
D. Okamura, R. Harakawa, and M. Iwahashi, “Lcnme: Label correction using network prediction based on memorization effects for cross-modal retrieval with noisy labels,” IEEE TCSVT , vol. 34, no. 1, pp. 590–602, 2024
2024
Closest in time.
Z. Dang, M. Luo, C. Jia, G. Dai, X. Chang, and J. Wang, “Noisy correspondence learning with self-reinforcing errors mitigation,” in AAAI , vol. 38, no. 2, 2024, pp. 1463–1471
2024
Closest in time.
Y. Qin, Y. Chen, D. Peng, X. Peng, J. T. Zhou, and P. Hu, “Noisy-correspondence learning for text-to-image person re-identification,” in IEEE CVPR , 2024, pp. 27 197–27 206
2024
Closest in time.
D. Shi, L. Zhu, J. Li, G. Dong, and H. Zhang, “Incomplete cross-modal retrieval with deep correlation transfer,” ACM TOMM , vol. 20, no. 5, pp. 1–21, 2024
2024
Closest in time.
H. Luo, Z. Zhang, and L. Nie, “Contrastive incomplete cross-modal hashing,” IEEE TKDE , pp. 1–12, 2024
2024
Closest in time.
Y.-W. Zhan, X. Luo, Z.-D. Chen, Y. Wang, Y. Wei, and X.-S. Xu, “Polish: Adaptive online cross-modal hashing for class incremental data,” in ACM WWW , 2024, pp. 4470–4478
2024
Closest in time.
F. Li, B. Wang, L. Zhu, J. Li, Z. Zhang, and X. Chang, “Cross-domain transfer hashing for efficient cross-modal retrieval,” IEEE TCSVT , pp. 1–14, 2024
2024
Closest in time.
R. Zuo, C. Zheng, F. Li, L. Zhu, and Z. Zhang, “Privacy-enhanced prototype-based federated cross-modal hashing for cross-modal retrieval,” ACM TOMM , 2024
2024
Closest in time.
T. Fu, Y.-W. Zhan, C.-Y. Zhang, X. Luo, Z.-D. Chen, Y. Wang, X. Yang, and X.-S. Xu, “Fedcafe: Federated cross-modal hashing with adaptive feature enhancement,” in ACM MM , 2024
2024
Closest in time.
Y. Xu, Y. Bin, J. Wei, Y. Yang, G. Wang, and H. T. Shen, “Align and retrieve: Composition and decomposition learning in image retrieval with text feedback,” IEEE TMM , pp. 1–13, 2024
2024
Closest in time.
Z. Zhang, X. Yuan, L. Zhu, J. Song, and L. Nie, “Badcm: Invisible backdoor attack against cross-modal learning,” IEEE TIP , vol. 33, pp. 2558–2571, 2024
2024
Closest in time.
T. Wang, F. Li, L. Zhu, J. Li, Z. Zhang, and H. T. Shen, “Invisible black-box backdoor attack against deep cross-modal hashing retrieval,” ACM TOIS , vol. 42, no. 4, pp. 1–27, 2024
2024
Closest in time.
M. Jin, W. Hu, L. Zhu, X. Wang, and R. Hong, “Based on spatial and temporal implicit semantic relational inference for cross-modal retrieval,” IEEE TCSVT , 2024
2024
Closest in time.
X. Yang, L. Zhu, X. Wang, and Y. Yang, “Dgl: Dynamic global-local prompt tuning for text-video retrieval,” in AAAI , vol. 38, no. 7, 2024, pp. 6540–6548
2024
Closest in time.
K. Tian, R. Zhao, Z. Xin, B. Lan, and X. Li, “Holistic features are almost sufficient for text-to-video retrieval,” in IEEE CVPR , 2024, pp. 17 138–17 147
2024
Closest in time.
J. Wang, G. Sun, P. Wang, D. Liu, S. Dianat, M. Rabbani, R. Rao, and Z. Tao, “Text is mass: Modeling as stochastic embedding for text-video retrieval,” in IEEE CVPR , 2024, pp. 16 551–16 560
2024
Closest in time.
D. Han, X. Cheng, N. Guo, X. Ye, B. Rainer, and P. Priller, “Momentum cross-modal contrastive learning for video moment retrieval,” IEEE TCSVT , vol. 34, no. 7, pp. 5977–5994, 2024
2024
Closest in time.
X. Fang, D. Liu, W. Fang, P. Zhou, Z. Xu, W. Xu, J. Chen, and R. Li, “Fewer steps, better performance: Efficient cross-modal clip trimming for video moment retrieval using language,” in AAAI , vol. 38, no. 2, 2024, pp. 1735–1743
2024
Closest in time.
D. Zhou, F. Lei, L. Li, Y. Zhou, and A. Yang, “Cross-modal interaction via reinforcement feedback for audio-lyrics retrieval,” IEEE/ACM TASLP , vol. 32, pp. 1248–1260, 2024
2024
Closest in time.
S. Doh, M. Lee, D. Jeong, and J. Nam, “Enriching music descriptions with a finetuned-llm and metadata for text-to-music retrieval,” in IEEE ICASSP , 2024, pp. 826–830
2024
Closest in time.
J. Huang, Y. Chen, S. Xiong, and X. Lu, “Cross-modal remote sensing image-audio retrieval with adaptive learning for aligning correlation,” IEEE TGRS , vol. 62, pp. 1–13, 2024
2024
Closest in time.
X. Qian, W. Xue, Q. Zhang, R. Tao, and H. Li, “Deep cross-modal retrieval between spatial image and acoustic speech,” IEEE TMM , vol. 26, pp. 4480–4489, 2024
2024
Closest in time.
J. Wang, A. Zheng, Y. Yan, R. He, and J. Tang, “Attribute-guided cross-modal interaction and enhancement for audio-visual matching,” IEEE TIFS , vol. 19, pp. 4986–4998, 2024
2024
Closest in time.
F. Zhang, X.-S. Hua, C. Chen, and X. Luo, “Fine-grained prototypical voting with heterogeneous mixup for semi-supervised 2d-3d cross-modal retrieval,” in IEEE CVPR , 2024, pp. 17 016–17 026
2024
Closest in time.
F. Zhang, H. Zhou, X.-S. Hua, C. Chen, and X. Luo, “Hope: A hierarchical perspective for semi-supervised 2d-3d cross-modal retrieval,” IEEE TPAMI , pp. 1–18, 2024
2024
Closest in time.
Y. Zhang, R. Hu, R. Li, Y. Qu, Y. Xie, and X. Li, “Cross-modal match for language conditioned 3d object grounding,” in AAAI , vol. 38, no. 7, 2024, pp. 7359–7367
2024
Closest in time.
2024
Closest in time.
H. Liu, J. Wu, F. Li, J. Jiang, and R. Hong, “Syrer: Synergistic relational reasoning for rgb-d cross-modal re-identification,” IEEE TMM , vol. 26, pp. 5600–5614, 2024
2024
Closest in time.
2024
Closest in time.
Z. Wang, Z. Gao, Y. Yang, G. Wang, C. Jiao, and H. T. Shen, “Geometric matching for cross-modal retrieval,” IEEE TNNLS , pp. 1–13, 2024
2024
Closest in time.
X. Huang, J. Liu, Z. Zhang, Y. Xie, Y. Tang, W. Zhang, and X. Cui, “Cross-modal recipe retrieval with fine-grained prompting alignment and evidential semantic consistency,” IEEE TMM , pp. 1–12, 2024
2024
Closest in time.
J. Zuo, H. Zhou, Y. Nie, F. Zhang, T. Guo, N. Sang, Y. Wang, and C. Gao, “Ufinebench: Towards text-based person retrieval with ultra-fine granularity,” in IEEE CVPR , 2024, pp. 22 010–22 019
2024
Closest in time.
L. Zhu, Y. Wang, Y. Hu, X. Su, and K. Fu, “Cross-modal contrastive learning with spatiotemporal context for correlation-aware multiscale remote sensing image retrieval,” IEEE TGRS , vol. 62, pp. 1–21, 2024
2024
Closest in time.
L. Zhang and X. Gao, “Transfer adaptation learning: A decade survey,” IEEE TNNLS , vol. 35, no. 1, pp. 23–44, 2024
2024
Closest in time.
T. Zhang and J. Wang, “Collaborative quantization for cross-modal similarity search,” in IEEE CVPR , 2016, pp. 2036–2045
2045
Closest in time.
G. Ding, Y. Guo, and J. Zhou, “Collective matrix factorization hashing for multimodal data,” in IEEE CVPR , 2014, pp. 2083–2090
2090
Closest in time.