2023

Teaching CLIP to Count to Ten

Paiss, Roni, Ephrat, Ariel, Tov, Omer et al.

Understand

Large vision-language models (VLMs), such as CLIP, learn rich joint image-text representations, facilitating advances in numerous downstream tasks, including zero-shot classification and text-to-image generation.

  • Nevertheless, existing VLMs exhibit a prominent well-documented limitation - they fail to encapsulate compositional concepts such as counting.
  • We introduce a simple yet effective method to improve the quantitative understanding of VLMs, while maintaining their overall performance on common benchmarks.
  • Specifically, we propose a new counting-contrastive loss used to finetune a pre-trained VLM in tandem with its original objective.

Reading the bibliography…