2024

CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling

Zhang, Jihai, Qu, Xiaoye, Zhu, Tong et al.

Understand

Contrastive Language-Image Pre-training (CLIP) has become a cornerstone in multimodal intelligence.

  • However, recent studies discovered that CLIP can only encode one aspect of the feature space, leading to substantial information loss and indistinctive features.
  • To mitigate this issue, this paper introduces a novel strategy that fine-tunes a series of complementary CLIP models and transforms them into a CLIP-MoE.
  • Specifically, we propose a model-agnostic Diversified Multiplet Upcycling (DMU) framework for CLIP.

Reading the bibliography…