Fetching the paper…

DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation · Around