Fetching the paper…

Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation · Around