Fetching the paper…

Contextualized Diffusion Models for Text-Guided Image and Video Generation · Around