2022

RoentGen: Vision-Language Foundation Model for Chest X-ray Generation

Chambon, Pierre, Bluethgen, Christian, Delbrouck, Jean-Benoit et al.

Understand

Multimodal models trained on large natural image-text pair datasets have exhibited astounding abilities in generating high-quality images.

  • Medical imaging data is fundamentally different to natural images, and the language used to succinctly capture relevant details in medical data uses a different, narrow but semantically rich, domain-specific vocabulary.
  • Not surprisingly, multi-modal models trained on natural image-text pairs do not tend to generalize well to the medical domain.
  • Developing generative imaging models faithfully representing medical concepts while providing compositional diversity could mitigate the existing paucity of high-quality, annotated medical imaging datasets.

Reading the bibliography…