Fetching the paper…

Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model · Around