2023

DocumentCLIP: Linking Figures and Main Body Text in Reflowed Documents

Liu, Fuxiao, Tan, Hao, Tensmeyer, Chris

Understand

Vision-language pretraining models have achieved great success in supporting multimedia applications by understanding the alignments between images and text.

  • While existing vision-language pretraining models primarily focus on understanding single image associated with a single piece of text, they often ignore the alignment at the intra-document level, consisting of multiple sentences with multiple images.
  • In this work, we propose DocumentCLIP, a salience-aware contrastive learning framework to enforce vision-language pretraining models to comprehend the interaction between images and longer text within documents.
  • Our model is beneficial for the real-world multimodal document understanding like news article, magazines, product descriptions, which contain linguistically and visually richer content.

Reading the bibliography…