2023

MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens

Zheng, Kaizhi, He, Xuehai, Wang, Xin Eric

Understand

The effectiveness of Multimodal Large Language Models (MLLMs) demonstrates a profound capability in multimodal understanding.

  • However, the simultaneous generation of images with coherent texts is still underdeveloped.
  • Addressing this, we introduce a novel interleaved vision-and-language generation method, centered around the concept of ``generative vokens".
  • These vokens serve as pivotal elements contributing to coherent image-text outputs.

Built on

Nothing clear enough to list yet.

Similar

Nothing clear enough to list yet.

Then

Nothing clear enough to list yet.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…