Fetching the paper…

Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation · Around