Fetching the paper…
Reading the bibliography…
Recently, multi-modal vision-language foundation models have gained significant attention in the medical field.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778 (2016)
2016
Earlier work this paper cites.
Sohn, K.: Improved deep metric learning with multi-class n-pair loss objective. Advances in neural information processing systems 29
2016
Earlier work this paper cites.
Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., Summers, R.M.: Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2097–2106 (2017)
2017
Earlier work this paper cites.
Chen, L., Bentley, P., Mori, K., Misawa, K., Fujiwara, M., Rueckert, D.: Self-supervised learning for medical image analysis using image context restoration. Medical image analysis 58
2019
Earlier work this paper cites.
Kenton, J.D.M.-W.C., Toutanova, L.K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of naacL-HLT, vol. 1, p. 2 (2019)
2019
Earlier work this paper cites.
Johnson, A.E., Pollard, T.J., Berkowitz, S.J., Greenbaum, N.R., Lungren, M.P., Deng, C.-y., Mark, R.G., Horng, S.: Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports. Scientific data 6
2019
Earlier work this paper cites.
Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., et al
2019
Earlier work this paper cites.
Shih, G., Wu, C.C., Halabi, S.S., Kohli, M.D., Prevedello, L.M., Cook, T.S., Sharma, A., Amorosa, J.K., Arteaga, V., Galperin-Aizenberg, M., et al
2019
Earlier work this paper cites.
Zawacki, A., Wu, C., Shih, G., Elliott, J., Fomitchev, M., Hussain, M., ParasLakhani, Culliton, P., Bao, S.: SIIM-ACR Pneumothorax Segmentation. https://kaggle.com/competitions/siim-acr-pneumothorax-segmentation (2019)
2019
Earlier work this paper cites.
Sutton, R.T., Pincock, D., Baumgart, D.C., Sadowski, D.C., Fedorak, R.N., Kroeker, K.I.: An overview of clinical decision support systems: benefits, risks, and strategies for success. NPJ digital medicine 3
2020
Earlier work this paper cites.
Zhou, H.-Y., Yu, S., Bian, C., Hu, Y., Ma, K., Zheng, Y.: Comparing to learn: Surpassing imagenet pretraining on radiographs by comparing image representations. In: Medical Image Computing and Computer Assisted Intervention–MICCAI 2020: 23rd International Conference, Lima, Peru, October 4–8, 2020, Proceedings, Part I 23, pp. 398–407 (2020). Springer
2020
Earlier work this paper cites.
Misra, I., Maaten, L.v.d.: Self-supervised learning of pretext-invariant representations. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6707–6717 (2020)
2020
Earlier work this paper cites.
Jaiswal, A., Babu, A.R., Zadeh, M.Z., Banerjee, D., Makedon, F.: A survey on contrastive self-supervised learning. Technologies 9
2020
Earlier work this paper cites.
Tang, H., Sun, N., Li, Y., Xia, H.: Deep learning segmentation model for automated detection of the opacity regions in the chest x-rays of the covid-19 positive patients and the application for disease severity. medRxiv, 2020–10 (2020)
2020
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al
2021
Earlier work this paper cites.
Huang, S.-C., Shen, L., Lungren, M.P., Yeung, S.: GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-efficient Medical Image Recognition. In: 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3922–3931 (2021). https://doi.org/10.1109/ICCV48922.2021.00391
2021
Earlier work this paper cites.
Zhou, Z., Sodha, V., Pang, J., Gotway, M.B., Liang, J.: Models genesis. Medical image analysis 67
2021
Earlier work this paper cites.
Haghighi, F., Taher, M.R.H., Zhou, Z., Gotway, M.B., Liang, J.: Transferable visual words: Exploiting the semantics of anatomical patterns for self-supervised learning. IEEE transactions on medical imaging 40
2021
Cited alongside, same era.
Acosta, J.N., Falcone, G.J., Rajpurkar, P., Topol, E.J.: Multimodal biomedical AI. Nature Medicine 28
2022
Cited alongside, same era.
Tiu, E., Talius, E., Patel, P., Langlotz, C.P., Ng, A.Y., Rajpurkar, P.: Expert-level detection of pathologies from unannotated chest x-ray images via self-supervised learning. Nature Biomedical Engineering 6
2022
Cited alongside, same era.
Zhou, H.-Y., Chen, X., Zhang, Y., Luo, R., Wang, L., Yu, Y.: Generalized radiograph representation learning via cross-supervision between images and free-text radiology reports. Nature Machine Intelligence 4
2022
Cited alongside, same era.
Wu, C., Zhang, X., Zhang, Y., Wang, Y., Xie, W.: Medklip: Medical knowledge enhanced language-image pre-training for x-ray diagnosis. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 21372–21383 (2023)
2023
Closest in time.
Zhou, Y., Chia, M.A., Wagner, S.K., Ayhan, M.S., Williamson, D.J., Struyven, R.R., Liu, T., Xu, M., Lozano, M.G., Woodward-Court, P., et al.: A foundation model for generalizable disease detection from retinal images. Nature, 1–8 (2023)
2023
Closest in time.
Zhou, H.-Y., et al
2023
Closest in time.
Zhang, X., Wu, C., Zhang, Y., Xie, W., Wang, Y.: Knowledge-enhanced visual-language pre-training on chest radiology images. Nature Communications 14
2023
Closest in time.
Zhou, H.-Y., Yu, Y., Wang, C., Zhang, S., Gao, Y., Pan, J., Shao, J., Lu, G., Zhang, K., Li, W.: A transformer-based representation-learning model with unified processing of multimodal input for clinical diagnostics. Nature Biomedical Engineering, 1–13 (2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., Girshick, R.: Masked Autoencoders Are Scalable Vision Learners. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 15979–15988 (2022). https://doi.org/10.1109/CVPR52688.2022.01553
2022
Cited alongside, same era.
Chen, Z., Du, Y., Hu, J., Liu, Y., Li, G., Wan, X., Chang, T.-H.: Multi-modal masked autoencoders for medical vision-and-language pre-training. In: International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 679–689 (2022). Springer
2022
Cited alongside, same era.
Boecking, B., Usuyama, N., Bannur, S., Castro, D.C., Schwaighofer, A., Hyland, S., Wetscherek, M., Naumann, T., Nori, A., Alvarez-Valle, J., et al
2022
Cited alongside, same era.
Müller, P., Kaissis, G., Zou, C., Rueckert, D.: Joint learning of localized representations from medical images and reports. In: European Conference on Computer Vision, pp. 685–701 (2022). Springer
2022
Cited alongside, same era.
Li, Y., Mao, H., Girshick, R., He, K.: Exploring plain vision transformer backbones for object detection. In: European Conference on Computer Vision, pp. 280–296 (2022). Springer
2022
Cited alongside, same era.
Albelwi, S.: Survey on self-supervised learning: auxiliary pretext tasks and contrastive learning methods in imaging. Entropy 24
2022
Cited alongside, same era.
Geng, X., Liu, H., Lee, L., Schuurmans, D., Levine, S., Abbeel, P.: Multimodal Masked Autoencoders Learn Transferable Representations. In: First Workshop on Pre-training: Perspectives, Pitfalls, and Paths Forward at ICML 2022 (2022)
2022
Cited alongside, same era.
Zhang, Y., Jiang, H., Miura, Y., Manning, C.D., Langlotz, C.P.: Contrastive learning of medical visual representations from paired images and text. In: Machine Learning for Healthcare Conference, pp. 2–25 (2022). PMLR
2022
Cited alongside, same era.
2023
Closest in time.
Huang, Z., Bianchi, F., Yuksekgonul, M., Montine, T.J., Zou, J.: A visual–language foundation model for pathology image analysis using medical twitter. Nature Medicine, 1–10 (2023)
2023
Closest in time.
Zhou, H.-Y., Lian, C., Wang, L., Yu, Y.: Advancing Radiograph Representation Learning with Masked Record Modeling. In: The Eleventh International Conference on Learning Representations (2023)
2023
Closest in time.
Bannur, S., Hyland, S., Liu, Q., Perez-Garcia, F., Ilse, M., Castro, D.C., Boecking, B., Sharma, H., Bouzid, K., Thieme, A., et al
2023
Closest in time.
Li, Y., Yang, B., Cheng, X., Zhu, Z., Li, H., Zou, Y.: Unify, align and refine: Multi-level semantic alignment for radiology report generation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2863–2874 (2023)
2023
Closest in time.
Liu, C., Cheng, S., Chen, C., Qiao, M., Zhang, W., Shah, A., Bai, W., Arcucci, R.: M-FLAG: Medical Vision-Language Pre-training with Frozen Language Models and Latent Space Geometry Optimization. In: International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 637–647 (2023)
2023
Closest in time.
Ma, D., Pang, J., Gotway, M.B., Liang, J.: Foundation Ark: Accruing and Reusing Knowledge for Superior and Robust Performance. In: International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 651–662 (2023). Springer
2023
Closest in time.
2024
Closest in time.
Liu, J., Zhou, H.-Y., Li, C., Huang, W., Yang, H., Liang, Y., Wang, S.: Mlip: Medical language-image pre-training with masked local representation learning. In: 2024 IEEE 21th International Symposium on Biomedical Imaging (ISBI) (2024)
2024
Closest in time.
Yang, H., Zhou, H.-Y., Li, C., Huang, W., Liu, J., Liang, Y., Wang, S.: Multimodal self-supervised learning for lesion localization. In: 2024 IEEE 21th International Symposium on Biomedical Imaging (ISBI) (2024)
2024
Closest in time.
Huang, W., Li, C., Zhou, H.-Y., Liu, J., Yang, H., Liang, Y., Shi, G., Zheng, H., Wang, S.: Enhancing representation in medical vision-language foundation models via multi-scale information extraction techniques. In: 2024 IEEE 21th International Symposium on Biomedical Imaging (ISBI) (2024)
2024
Closest in time.
Wan, Z., Liu, C., Zhang, M., Fu, J., Wang, B., Cheng, S., Ma, L., Quilodrán-Casas, C., Arcucci, R.: Med-unic: Unifying cross-lingual medical vision-language pre-training by diminishing bias. Advances in Neural Information Processing Systems 36
2024
Closest in time.