2020

Crisscrossed Captions: Extended Intramodal and Intermodal Semantic Similarity Judgments for MS-COCO

Parekh, Zarana, Baldridge, Jason, Cer, Daniel et al.

Understand

By supporting multi-modal retrieval training and evaluation, image captioning datasets have spurred remarkable progress on representation learning.

  • Unfortunately, datasets have limited cross-modal associations: images are not paired with other images, captions are only paired with other captions of the same image, there are no negative associations and there are missing positive cross-modal associations.
  • This undermines research into how inter-modality learning impacts intra-modality tasks.
  • We address this gap with Crisscrossed Captions (CxC), an extension of the MS-COCO dataset with human semantic similarity judgments for 267,095 intra- and inter-modality pairs.

Reading the bibliography…