Fetching the paper…
Reading the bibliography…
Medical Image Grounding (MIG), which involves localizing specific regions in medical images based on textual descriptions, requires models to not only perceive regions but also deduce spatial relationships of these regions.
2017
Earlier work this paper cites.
Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., Summers, R.M.: Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 2097–2106 (2017)
2017
Earlier work this paper cites.
Johnson, A.E., Pollard, T.J., Berkowitz, S.J., Greenbaum, N.R., Lungren, M.P., Deng, C.y., Mark, R.G., Horng, S.: Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports. Scientific data 6
2019
Earlier work this paper cites.
Yu, L., Zhang, Z., Li, X., Ren, H., Zhao, W., Xing, L.: Metal artifact reduction in 2d ct images with self-supervised cross-domain learning. Physics in Medicine & Biology 66
2021
Earlier work this paper cites.
Boecking, B., Usuyama, N., Bannur, S., Castro, D.C., Schwaighofer, A., Hyland, S., Wetscherek, M., Naumann, T., Nori, A., Alvarez-Valle, J., et al.: Making the most of text semantics to improve biomedical vision–language processing. In: European conference on computer vision. pp. 1–21. Springer (2022)
2022
Earlier work this paper cites.
Wang, Z., Wu, Z., Agarwal, D., Sun, J.: Medclip: Contrastive learning from unpaired medical images and text. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing. vol. 2022, p. 3876 (2022)
2022
Earlier work this paper cites.
Chen, Z., Zhou, Y., Tran, A., Zhao, J., Wan, L., Ooi, G.S.K., Cheng, L.T.E., Thng, C.H., Xu, X., Liu, Y., et al.: Medical phrase grounding with region-phrase context contrastive alignment. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. pp. 371–381. Springer (2023)
2023
Earlier work this paper cites.
Li, Z., Li, Y., Li, Q., Wang, P., Guo, D., Lu, L., Jin, D., Zhang, Y., Hong, Q.: Lvit: language meets vision transformer in medical image segmentation. IEEE transactions on medical imaging 43
2023
Earlier work this paper cites.
Tian, X., Ng, W.W., Xu, H.: Deep incremental hashing for semantic image retrieval with concept drift. IEEE Transactions on Big Data 9
2023
Earlier work this paper cites.
Wasserthal, J., Breit, H.C., Meyer, M.T., Pradella, M., Hinck, D., Sauter, A.W., Heye, T., Boll, D.T., Cyriac, J., Yang, S., et al.: Totalsegmentator: robust segmentation of 104 anatomic structures in ct images. Radiology: Artificial Intelligence 5
2023
Earlier work this paper cites.
Zhong, Y., Xu, M., Liang, K., Chen, K., Wu, M.: Ariadne’s thread: Using text prompts to improve segmentation of infected areas from chest x-ray images. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. pp. 724–733. Springer (2023)
2023
Cited alongside, same era.
2024
Cited alongside, same era.
Chen, Y., Wei, M., Zheng, Z., Hu, J., Shi, Y., Xiong, S., Zhu, X.X., Mou, L.: Causalclipseg: Unlocking clip’s potential in referring medical image segmentation with causal intervention. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. pp. 77–87. Springer (2024)
2024
Cited alongside, same era.
2024
Later among the works it cites.
Wu, H., Yang, Y., Xu, H., Wang, W., Zhou, J., Zhu, L.: Rainmamba: Enhanced locality learning with state space models for video deraining. In: Proceedings of the 32nd ACM International Conference on Multimedia. pp. 7881–7890 (2024)
2024
Later among the works it cites.
Xu, H., Yang, Y., Aviles-Rivero, A.I., Yang, G., Qin, J., Zhu, L.: Lgrnet: Local-global reciprocal network for uterine fibroid segmentation in ultrasound videos. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. pp. 667–677. Springer (2024)
2024
Later among the works it cites.
Xu, H., Zhu, L.: Online agglomerative pooling for scalable self-supervised universal segmentation (2024), https://openreview.net/forum?id=d32d9fE5lG
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
Huang, X., Huang, H., Shen, L., Yang, Y., Shang, F., Liu, J., Liu, J.: A refer-and-ground multimodal large language model for biomedicine. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. pp. 399–409. Springer (2024)
2024
Cited alongside, same era.
Huang, X., Li, H., Cao, M., Chen, L., You, C., An, D.: Cross-modal conditioned reconstruction for language-guided medical image segmentation. IEEE Transactions on Medical Imaging (2024)
2024
Cited alongside, same era.
Müller, P., Kaissis, G., Rueckert, D.: Chex: Interactive localization and region description in chest x-rays. In: European Conference on Computer Vision. pp. 92–111. Springer (2024)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Wang, H., Wang, W., Zhou, H., Xu, H., Wu, S., Zhu, L.: Language-driven interactive shadow detection. In: Proceedings of the 32nd ACM International Conference on Multimedia. pp. 5527–5536 (2024)
2024
Cited alongside, same era.
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.
Chen, Y., Wang, G., Ji, Y., Li, Y., Ye, J., Li, T., Hu, M., Yu, R., Qiao, Y., He, J.: Slidechat: A large vision-language assistant for whole-slide pathology image understanding. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 5134–5143 (2025)
2025
Closest in time.
2025
Closest in time.
Li, Y., Wang, H., Duan, Y., Zhang, J., Li, X.: A closer look at the explainability of contrastive language-image pre-training. Pattern Recognition p. 111409 (2025)
2025
Closest in time.