2024

MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Tian, Changyao, Zhu, Xizhou, Xiong, Yuwen et al.

Understand

Developing generative models for interleaved image-text data has both research and practical value.

  • It requires models to understand the interleaved sequences and subsequently generate images and text.
  • However, existing attempts are limited by the issue that the fixed number of visual tokens cannot efficiently capture image details, which is particularly problematic in the multi-image scenarios.
  • To address this, this paper presents MM-Interleaved, an end-to-end generative model for interleaved image-text data.

Reading the bibliography…