Fetching the paper…
Reading the bibliography…
The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model.
Collecting highly parallel data for paraphrase evaluation
D. Chen and W. Dolan · 2011
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Generative adversarial nets
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow
Z. Teed and J. Deng · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
M. Bain, A. Nagrani, G. Varol, and A. Zisserman · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki · 2021
Earlier work this paper cites.
Merlot: Multimodal neural script knowledge models
R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi · 2021
Earlier work this paper cites.
aesthetic-predictor
L. AI · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Cited alongside, same era.
Advancing high-resolution video-language representation with large-scale video transcriptions
H. Xue, T. Hang, Y. Zeng, Y. Sun, B. Liu, H. Yang, J. Fu, and B. Guo · 2022
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, et al · 2023
Cited alongside, same era.
Paddleocr
PaddlePaddle · 2023
Cited alongside, same era.
Sora | openai, 2024
C. Ng, D. Schnurr, E. Luhman, J. Taylor, L. Jing, N. Summers, R. Wang, R. Sahai, R. O’Rourke, T. Luhman, W. DePue, and Y. Guo · 2024
Later among the works it cites.
SDXL: Improving latent diffusion models for high-resolution image synthesis
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach · 2024
Later among the works it cites.
Vidgen-1m: A large-scale dataset for text-to-video generation
Z. Tan, X. Yang, L. Qin, and H. Li · 2024
Later among the works it cites.
Lvd-2m: A long-take video dataset with temporally dense captions
T. Xiong, Y. Wang, D. Zhou, Z. Lin, J. Feng, and X. Liu · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
J. Z. Wu, Y. Ge, X. Wang, S. W. Lei, Y. Gu, Y. Shi, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou · 2023
Cited alongside, same era.
Pyscenedetect: Python-based video scene detector, March 2024
B. Castellano · 2024
Cited alongside, same era.
Autoregressive video generation without vector quantization
H. Deng, T. Pan, H. Diao, Z. Luo, Y. Cui, H. Lu, S. Shan, Y. Qi, and X. Wang · 2024
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Y. Guo, C. Yang, A. Rao, Z. Liang, Y. Wang, Y. Qiao, M. Agrawala, D. Lin, and B. Dai · 2024
Cited alongside, same era.
Vbench: Comprehensive benchmark suite for video generative models
Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, et al · 2024
Cited alongside, same era.
Miradata: A large-scale video dataset with long durations and structured captions
X. Ju, Y. Gao, Z. Zhang, Z. Yuan, X. Wang, A. Zeng, Y. Xiong, Q. Xu, and Y. Shan · 2024
Cited alongside, same era.
Hunyuanvideo: A systematic framework for large video generative models
W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al · 2024
Cited alongside, same era.
Later among the works it cites.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al · 2025
Closest in time.
Step-video-t2v technical report: The practice, challenges, and future of video foundation model
G. Ma, H. Huang, K. Yan, L. Chen, N. Duan, S. Yin, C. Wan, R. Ming, X. Song, X. Chen, et al · 2025
Closest in time.
Openvid-1m: A large-scale high-quality dataset for text-to-video generation
K. Nan, R. Xie, P. Zhou, T. Fan, Z. Yang, Z. Chen, X. Li, J. Yang, and Y. Tai · 2025
Closest in time.
Magi-1: Autoregressive video generation at scale, 2025
Sand-AI · 2025
Closest in time.
Streamlit
Streamlit · 2025
Closest in time.
Qwen3, April 2025
Q. Team · 2025
Closest in time.
Wan: Open and advanced large-scale video generative models
A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, J. Zeng, J. Wang, J. Zhang, J. Zhou, J. Wang, J. Chen, K. Zhu, K. Zhao, K. Yan, L. Huang, M. Feng, N. Zhang, P. Li, P. Wu, R. Chu, R. Feng, S. Zhang, S. Sun, T. Fang, T. Wang, T. Gui, T. Weng, T. Shen, W. Lin, W. Wang, W. Wang, W. Zhou, W. Wang, W. Shen, W. Yu, X. Shi, X. Huang, X. Xu, Y. Kou, Y. Lv, Y. Li, Y. Liu, Y. Wang, Y. Zhang, Y. Huang, Y. Li, Y. Wu, Y. Liu, Y. Pan, Y. Zheng, Y. Hong, Y. Shi, Y. Feng, Z. Jiang, Z. Han, Z.-F. Wu, and Z. Liu · 2025
Closest in time.
Videoufo: A million-scale user-focused dataset for text-to-video generation
W. Wang and Y. Yang · 2025
Closest in time.