2023

DreamLLM: Synergistic Multimodal Comprehension and Creation

Dong, Runpei, Han, Chunrui, Peng, Yuang et al.

Understand

This paper presents DreamLLM, a learning framework that first achieves versatile Multimodal Large Language Models (MLLMs) empowered with frequently overlooked synergy between multimodal comprehension and creation.

  • DreamLLM operates on two fundamental principles.
  • The first focuses on the generative modeling of both language and image posteriors by direct sampling in the raw multimodal space.
  • This approach circumvents the limitations and information loss inherent to external feature extractors like CLIP, and a more thorough multimodal understanding is obtained.

Reading the bibliography…