Fetching the paper…
Reading the bibliography…
Recent advancements in video generation have enabled models to synthesize high-quality, minute-long videos.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2021
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding, 2021
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Earlier work this paper cites.
Godiva: Generating open-domain videos from natural descriptions
Wu, C., Huang, L., Zhang, Q., Li, B., Ji, L., Yang, F., Sapiro, G., and Duan, N · 2021
Earlier work this paper cites.
Videogpt: Video generation using vq-vae and transformers
Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A · 2021
Earlier work this paper cites.
Long video generation with time-agnostic vqgan and time-sensitive transformer
Ge, S., Hayes, T., Yang, H., Yin, X., Pang, G., Jacobs, D., Huang, J.-B., and Parikh, D · 2022
Earlier work this paper cites.
Latent video diffusion models for high-fidelity long video generation
He, Y., Yang, T., Zhang, Y., Shan, Y., and Chen, Q · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., and Salimans, T · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Earlier work this paper cites.
Nuwa-infinity: Autoregressive over autoregressive generation for infinite visual synthesis
Liang, J., Wu, C., Hu, X., Gan, Z., Wang, J., Wang, L., Liu, Z., Fang, Y., and Duan, N · 2022
Earlier work this paper cites.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al · 2022
Earlier work this paper cites.
Nüwa: Visual synthesis pre-training for neural visual world creation
Wu, C., Liang, J., Ji, L., Yang, F., Fang, Y., Jiang, D., and Duan, N · 2022
Earlier work this paper cites.
Egsde: Unpaired image-to-image translation via energy-guided stochastic differential equations
Zhao, M., Bao, F., Li, C., and Zhu, J · 2022
Earlier work this paper cites.
All are worth words: A vit backbone for diffusion models
Bao, F., Nie, S., Xue, K., Cao, Y., Li, C., Su, H., and Zhu, J · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., Jampani, V., and Rombach, R · 2023
Cited alongside, same era.
NTK-Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation., 2023
bloc97 · 2023
Cited alongside, same era.
Videopoet: A large language model for zero-shot video generation
Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Schindler, G., Hornung, R., Birodkar, V., Yan, J., Chiu, M.-C., et al · 2023
Cited alongside, same era.
Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning
Lin, H., Zala, A., Cho, J., and Bansal, M · 2023
FIFO-diffusion: Generating infinite videos from text without training
Kim, J., Kang, J., Choi, J., and Han, B · 2024
Later among the works it cites.
Hunyuanvideo: A systematic framework for large video generative models
Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., et al · 2024
Later among the works it cites.
Arlon: Boosting diffusion transformers with autoregressive models for long video generation
Li, Z., Hu, S., Liu, S., Zhou, L., Choi, J., Meng, L., Guo, X., Li, J., Ling, H., and Wei, F · 2024
Later among the works it cites.
Open-sora plan: Open-source large video generation model
Lin, B., Ge, Y., Cheng, X., Li, Z., Zhu, B., Wang, S., He, X., Ye, Y., Yuan, S., Chen, L., et al · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models
Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A., Vyas, A., Shi, B., Ma, C.-Y., Chuang, C.-Y., Yan, D., Choudhary, D., Wang, D., Sethi, G., Pang, G., Ma, H., Misra, I., Hou, J., Wang, J., Jagadeesh, K., Li, K., Zhang, L., Singh, M., Williamson, M., Le, M., Yu, M., Singh, M. K., Zhang, P., Vajda, P., Duval, Q., Girdhar, R., Sumbaly, R., Rambhatla, S. S., Tsai, S., Azadi, S., Datta, S., Chen, S., Bell, S., Ramaswamy, S., Sheynin, S., Bhattacharya, S., Motwani, S., Xu, T., Li, T., Hou, T., Hsu, W.-N., Yin, X., Dai, X., Taigman, Y., Luo, Y., Liu, Y.-C., Wu, Y.-C., Zhao, Y., Kirstain, Y., He, Z., He, Z., Pumarola, A., Thabet, A., Sanakoyeu, A., Mallya, A., Guo, B., Araya, B., Kerr, B., Wood, C., Liu, C., Peng, C., Vengertsev, D., Schonfeld, E., Blanchard, E., Juefei-Xu, F., Nord, F., Liang, J., Hoffman, J., Kohler, J., Fire, K., Sivakumar, K., Chen, L., Yu, L., Gao, L., Georgopoulos, M., Moritz, R., Sampson, S. K., Li, S., Parmeggiani, S., Fine, S., Fowler, T., Petrovic, V., and Du, Y · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Yarn: Efficient context window extension of large language models
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
Cited alongside, same era.
Freenoise: Tuning-free longer video diffusion via noise rescheduling, 2023
Qiu, H., Xia, M., Zhang, Y., He, Y., Wang, X., Shan, Y., and Liu, Z · 2023
Cited alongside, same era.
Gen-l-video: Multi-text to long video generation via temporal co-denoising
Wang, F.-Y., Chen, W., Song, G., Ye, H.-J., Liu, Y., and Li, H · 2023
Cited alongside, same era.
Dynamicrafter: Animating open-domain images with video diffusion priors
Xing, J., Xia, M., Zhang, Y., Chen, H., Wang, X., Wong, T.-T., and Shan, Y · 2023
Cited alongside, same era.
Controlvideo: Adding conditional control for one shot text-to-video editing
Zhao, M., Wang, R., Bao, F., Li, C., and Zhu, J · 2023
Cited alongside, same era.
Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models
Bao, F., Xiang, C., Yue, G., He, G., Zhu, H., Zheng, K., Zhao, M., Liu, S., Wang, Y., and Zhu, J · 2024
Cited alongside, same era.
Later among the works it cites.
Generative multimodal models are in-context learners
Sun, Q., Cui, Y., Zhang, X., Zhang, F., Yu, Q., Wang, Y., Rao, Y., Liu, J., Huang, T., and Wang, X · 2024
Later among the works it cites.
Vila-u: a unified foundation model integrating visual understanding and generation
Wu, Y., Zhang, Z., Chen, J., Tang, H., Li, D., Fang, Y., Zhu, L., Xie, E., Yin, H., Yi, L., et al · 2024
Later among the works it cites.
Long video diffusion generation with segmented cross-attention and content-rich video data curation
Yan, X., Cai, Y., Wang, Q., Zhou, Y., Huang, W., and Yang, H · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., et al · 2024
Later among the works it cites.
From slow bidirectional to fast causal video generators
Yin, T., Zhang, Q., Zhang, R., Freeman, W. T., Durand, F., Shechtman, E., and Huang, X · 2024
Later among the works it cites.
Identifying and solving conditional image leakage in image-to-video diffusion model
Zhao, M., Zhu, H., Xiang, C., Zheng, K., Li, C., and Zhu, J · 2024
Later among the works it cites.
Open-sora: Democratizing efficient video production for all
Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., and You, Y · 2024
Later among the works it cites.
Allegro: Open the black box of commercial-level video generation model
Zhou, Y., Wang, Q., Cai, Y., and Yang, H · 2024
Later among the works it cites.
Lumina-next: Making lumina-t2x stronger and faster with next-dit
Zhuo, L., Du, R., Xiao, H., Li, Y., Liu, D., Huang, R., Liu, W., Zhao, L., Wang, F.-Y., Ma, Z., et al · 2024
Later among the works it cites.
Ouroboros-diffusion: Exploring consistent content generation in tuning-free long video diffusion
Chen, J., Long, F., An, J., Qiu, Z., Yao, T., Luo, J., and Mei, T · 2025
Closest in time.
Cosmos world foundation model platform for physical ai
NVIDIA, :, Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., Dworakowski, D., Fan, J., Fenzi, M., Ferroni, F., Fidler, S., Fox, D., Ge, S., Ge, Y., Gu, J., Gururani, S., He, E., Huang, J., Huffman, J., Jannaty, P., Jin, J., Kim, S. W., Klár, G., Lam, G., Lan, S., Leal-Taixe, L., Li, A., Li, Z., Lin, C.-H., Lin, T.-Y., Ling, H., Liu, M.-Y., Liu, X., Luo, A., Ma, Q., Mao, H., Mo, K., Mousavian, A., Nah, S., Niverty, S., Page, D., Paschalidou, D., Patel, Z., Pavao, L., Ramezanali, M., Reda, F., Ren, X., Sabavat, V. R. N., Schmerling, E., Shi, S., Stefaniak, B., Tang, S., Tchapmi, L., Tredak, P., Tseng, W.-C., Varghese, J., Wang, H., Wang, H., Wang, H., Wang, T.-C., Wei, F., Wei, X., Wu, J. Z., Xu, J., Yang, W., Yen-Chen, L., Zeng, X., Zeng, Y., Zhang, J., Zhang, Q., Zhang, Y., Zhao, Q., and Zolkowski, A · 2025
Closest in time.
Spargeattn: Accurate sparse attention accelerating any model inference
Zhang, J., Xiang, C., Huang, H., Wei, J., Xi, H., Zhu, J., and Chen, J · 2025
Closest in time.