2025

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

Mistretta, Marco, Baldrati, Alberto, Agnolucci, Lorenzo et al.

Understand

Pre-trained multi-modal Vision-Language Models like CLIP are widely used off-the-shelf for a variety of applications.

  • In this paper, we show that the common practice of individually exploiting the text or image encoders of these powerful multi-modal models is highly suboptimal for intra-modal tasks like image-to-image retrieval.
  • We argue that this is inherently due to the CLIP-style inter-modal contrastive loss that does not enforce any intra-modal constraints, leading to what we call intra-modal misalignment.
  • To demonstrate this, we leverage two optimization-based modality inversion techniques that map representations from their input modality to the complementary one without any need for auxiliary data or additional trained adapters.

Reading the bibliography…