2022

Multimodal Masked Autoencoders Learn Transferable Representations

Geng, Xinyang, Liu, Hao, Lee, Lisa et al.

Understand

Building scalable models to learn from diverse, multimodal data remains an open challenge.

  • For vision-language data, the dominant approaches are based on contrastive learning objectives that train a separate encoder for each modality.
  • While effective, contrastive learning approaches introduce sampling bias depending on the data augmentations used, which can degrade performance on downstream tasks.
  • Moreover, these methods are limited to paired image-text data, and cannot leverage widely-available unpaired data.

Reading the bibliography…