2022

Adapting Pretrained Vision-Language Foundational Models to Medical Imaging Domains

Chambon, Pierre, Bluethgen, Christian, Langlotz, Curtis P. et al.

Understand

Multi-modal foundation models are typically trained on millions of pairs of natural images and text captions, frequently obtained through web-crawling approaches.

  • Although such models depict excellent generative capabilities, they do not typically generalize well to specific domains such as medical images that have fundamentally shifted distributions compared to natural images.
  • Building generative models for medical images that faithfully depict clinical context may help alleviate the paucity of healthcare datasets.
  • Thus, in this study, we seek to research and expand the representational capabilities of large pretrained foundation models to medical concepts, specifically for leveraging the Stable Diffusion model to generate domain specific images found in medical imaging.

Reading the bibliography…