2022

Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts

Mustafa, Basil, Riquelme, Carlos, Puigcerver, Joan et al.

Understand

Large sparsely-activated models have obtained excellent performance in multiple domains.

  • However, such models are typically trained on a single modality at a time.
  • We present the Language-Image MoE, LIMoE, a sparse mixture of experts model capable of multimodal learning.
  • LIMoE accepts both images and text simultaneously, while being trained using a contrastive loss.

Reading the bibliography…