Fetching the paper…
Reading the bibliography…
Cross-modal medical image-report retrieval task plays a significant role in clinical diagnosis and various medical generative tasks.
Chen Y, Wang L, Wang W, et al. Continuum regression for cross-modal multimedia retrieval[C]. 2012 19th IEEE International Conference on Image Processing, 2012, 1949-1952
1952
Earlier work this paper cites.
Papineni K, Roukos S, Ward T, et al. Bleu: a method for automatic evaluation of machine translation[C]. Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 2002, 311-318
2002
Earlier work this paper cites.
Hardoon D. R, Szedmak S, Shawe-Taylor J. Canonical correlation analysis: An overview with application to learning methods[J]. Neural computation, 2004, 16(12): 2639-2664
2004
Earlier work this paper cites.
Lin C.-Y. Rouge: A package for automatic evaluation of summaries[C]. Text summarization branches out, 2004, 74-81
2004
Earlier work this paper cites.
Banerjee S, Lavie A. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments[C]. Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, 2005, 65-72
2005
Earlier work this paper cites.
Bird S. Nltk: the natural language toolkit[C]. Proceedings of the COLING/ACL 2006 Interactive Presentation Sessions, 2006, 69-72
2006
Earlier work this paper cites.
Maaten L. Vder , Hinton G. Visualizing data using t-sne.[J]. Journal of machine learning research, 2008, 9(11)
2008
Earlier work this paper cites.
Rasiwasia N, Pereira J. C, Coviello E, et al. A new approach to cross-modal multimedia retrieval[C]. Proceedings of the 18th ACM international conference on Multimedia, 2010, 251-260
2010
Earlier work this paper cites.
Schuster M, Nakajima K. Japanese and korean voice search[C]. 2012 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2012, 5149-5152
2012
Earlier work this paper cites.
Andrew G, Arora R, Bilmes J, et al. Deep canonical correlation analysis[C]. International conference on machine learning, 2013, 1247-1255
2013
Earlier work this paper cites.
Frome A, Corrado G. S, Shlens J, et al. Devise: A deep visual-semantic embedding model[J]. Advances in neural information processing systems, 2013, 26
2013
Earlier work this paper cites.
Gong Y, Wang L, Hodosh M, et al. Improving image-sentence embeddings using large weakly annotated photo collections[C]. Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part IV 13, 2014, 529-545
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Vedantam R, Zitnick C. L, Parikh D. Cider: Consensus-based image description evaluation[C]. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015,
2015
Earlier work this paper cites.
Zhang L, Ma B, Li G, et al. Pl-ranking: A novel ranking method for cross-modal retrieval[C]. Proceedings of the 24th ACM international conference on Multimedia, 2016, 1355-1364
2016
Earlier work this paper cites.
Faghri F, Fleet D. J, Kiros J. R, et al. Vse++: Improving visual-semantic embeddings with hard negatives[J]. , 2018
2018
Earlier work this paper cites.
Lee K.-H, Chen X, Hua G, et al. Stacked cross attention for image-text matching[C]. Proceedings of the European conference on computer vision (ECCV), 2018, 201-216
2018
Earlier work this paper cites.
Oord, A., Li, Y. & Vinyals, O. Representation learning with contrastive predictive coding
2018
Cited alongside, same era.
Wang Z, Liu X, Li H, et al. Camp: Cross-modal adaptive message passing for text-image retrieval[C]. Proceedings of the IEEE/CVF international conference on computer vision, 2019, 5764-5773
2019
Cited alongside, same era.
Johnson A. E, Pollard T. J, Berkowitz S. J, et al. Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports[J]. Scientific data, 2019, 6(1): 317
2019
Cited alongside, same era.
Alsentzer E, Murphy J. R, Boag W, et al. Publicly available clinical bert embeddings[J]. NAACL HLT 2019, 2019, 72
2019
Cited alongside, same era.
Zhang Q, Lei Z, Zhang Z, et al. Context-aware attention network for image-text retrieval[C]. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, 3536-3545
Ji Y, Tu R, Jiang J, et al. Seeing what you miss: Vision-language pre-training with semantic completion learning[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, 6789-6798
2023
Closest in time.
Dong X, Bao J, Zheng Y, et al. Maskclip: Masked self-distillation advances contrastive language-image pretraining[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, 10995-11005
2023
Closest in time.
Liu C, Zhang Y, Wang H, et al. Efficient token-guided image-text retrieval with consistent multimodal contrastive training[J]. IEEE Transactions on Image Processing, 2023
2023
Closest in time.
Kim D, Kim N, Kwak S. Improving cross-modal retrieval with set of diverse embeddings[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, 23422-23431
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Chun S, Oh S. J, Rezende R. S. D, et al. Probabilistic embeddings for cross-modal retrieval[C]. Conference on Computer Vision and Pattern Recognition (CVPR), 2021,
2021
Cited alongside, same era.
Radford A, Kim J. W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C]. International conference on machine learning, 2021, 8748-8763
2021
Cited alongside, same era.
Huang S.-C, Shen L, Lungren M. P, et al. Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition[C]. Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, 3942-3951
2021
Cited alongside, same era.
Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale[J]. ICLR, 2021
2021
Cited alongside, same era.
Liu Z, Chen F, Xu J, et al. Image-text retrieval with cross-modal semantic importance consistency[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022
2022
Cited alongside, same era.
Lu H, Fei N, Huo Y, et al. Cots: Collaborative two-stream vision-language pre-training model for cross-modal retrieval[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, 15692-15701
2022
Cited alongside, same era.
Moon J. H, Lee H, Shin W, et al. Multi-modal understanding and generation for medical images and text via vision-language pre-training[J]. IEEE Journal of Biomedical and Health Informatics, 2022, 26(12): 6070-6080
2022
Cited alongside, same era.
2023
Closest in time.
Wang X, Li L, Li Z, et al. Agree: Aligning cross-modal entities for image-text retrieval upon vision-language pre-trained models[C]. Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining, 2023, 456-464
2023
Closest in time.
Dawidowicz G, Hirsch E, Tal A. Limitr: Leveraging local information for medical image-text representation[C]. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, 21165-21173
2023
Closest in time.
You K, Gu J, Ham J, et al. Cxr-clip: Toward large scale chest x-ray language-image pre-training[C]. International Conference on Medical Image Computing and Computer-Assisted Intervention, 2023, 101-111
2023
Closest in time.
2023
Closest in time.
Chen C, Zhong A, Wu D, et al. Contrastive masked image-text modeling for medical visual representation learning[C]. International Conference on Medical Image Computing and Computer-Assisted Intervention, 2023, 493-503
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Li Y, Fan H, Hu R, et al. Scaling language-image pre-training via masking[C]. CVPR, 2023,
2023
Closest in time.
Chen, X., He, Y., Xue, C., Ge, R., Li, S. & Yang, G. Knowledge Boosting: Rethinking Medical Contrastive Vision-Language Pre-training
2023
Closest in time.
Liu, B., Lu, D., Wei, D., Wu, X., Wang, Y., Zhang, Y. & Zheng, Y. Improving Medical Vision-Language Contrastive Pretraining with Semantics-aware Triage
2023
Closest in time.