Fetching the paper…
Reading the bibliography…
Text-to-video diffusion models have shown remarkable progress in generating coherent video clips from textual descriptions.
A threshold selection method from gray-level histograms
Otsu, N · 1979
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
Perazzi, F., Pont-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., and Sorkine-Hornung, A · 2016
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Bain, M., Nagrani, A., Varol, G., and Zisserman, A · 2021
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Earlier work this paper cites.
Prompt-based multi-modal image segmentation
Lüddecke, T. and Ecker, A. S · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion, 2022
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2022
Earlier work this paper cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K · 2022
Earlier work this paper cites.
Cross-image attention for zero-shot appearance transfer, 2023
Alaluf, Y., Garibi, D., Patashnik, O., Averbuch-Elor, H., and Cohen-Or, D · 2023
Earlier work this paper cites.
The chosen one: Consistent characters in text-to-image diffusion models
Avrahami, O., Hertz, A., Vinker, Y., Arar, M., Fruchter, S., Fried, O., Cohen-Or, D., and Lischinski, D · 2023
Earlier work this paper cites.
Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing
Cao, M., Wang, X., Qi, Z., Shan, Y., Qie, X., and Zheng, Y · 2023
Earlier work this paper cites.
Pix2video: Video editing using image diffusion
Ceylan, D., Huang, C.-H. P., and Mitra, N. J · 2023
Earlier work this paper cites.
Magicdance: Realistic human dance video generation with motions & facial expressions transfer
Chang, D., Shi, Y., Gao, Q., Fu, J., Xu, H., Song, G., Yan, Q., Yang, X., and Soleymani, M · 2023
Earlier work this paper cites.
Flatten: optical flow-guided attention for consistent text-to-video editing
Cong, Y., Xu, M., Simon, C., Chen, S., Ren, J., Xie, Y., Perez-Rua, J.-M., Rosenhahn, B., Xiang, T., and He, S · 2023
Earlier work this paper cites.
Improved visual story generation with adaptive context modeling
Feng, Z., Ren, Y., Yu, X., Feng, X., Tang, D., Shi, S., and Qin, B · 2023
Earlier work this paper cites.
Dreamsim: Learning new dimensions of human visual similarity using synthetic data
Fu, S., Tamir, N. Y., Sundaram, S., Chai, L., Zhang, R., Dekel, T., and Isola, P · 2023
Earlier work this paper cites.
Encoder-based domain tuning for fast personalization of text-to-image models
Gal, R., Arar, M., Atzmon, Y., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2023
Earlier work this paper cites.
Tokenflow: Consistent diffusion features for consistent video editing
Geyer, M., Bar-Tal, O., Bagon, S., and Dekel, T · 2023
Earlier work this paper cites.
Style aligned image generation via shared attention
Hertz, A., Voynov, A., Fruchter, S., and Cohen-Or, D · 2023
Earlier work this paper cites.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Hu, L., Gao, X., Zhang, P., Sun, K., Zhang, B., and Bo, L · 2023
Cited alongside, same era.
Zero-shot generation of coherent storybook from plain text story using diffusion models
Jeong, H., Kwon, G., and Ye, J. C · 2023
Cited alongside, same era.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
Khachatryan, L., Movsisyan, A., Tadevosyan, V., Henschel, R., Wang, Z., Navasardyan, S., and Shi, H · 2023
Cited alongside, same era.
Intelligent grimm–open-ended visual storytelling via latent diffusion models
Liu, C., Wu, H., Zhong, Y., Zhang, X., and Xie, W · 2023
Cited alongside, same era.
Refdrop: Controllable consistency in image or video generation via reference feature guidance
Fan, J., Xue, H., Zhang, Q., and Chen, Y · 2024
Closest in time.
Lcm-lookahead for encoder-based text-to-image personalization, 2024
Gal, R., Lichter, O., Richardson, E., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2024
Closest in time.
Ltx-video: Realtime video latent diffusion
HaCohen, Y., Chiprut, N., Brazowski, B., Shalem, D., Moshe, D., Richardson, E., Levin, E., Shiran, G., Zabari, N., Gordon, O., Panet, P., Weissbuch, S., Kulikov, V., Bitterman, Y., Melumian, Z., and Bibi, O · 2024
Closest in time.
VBench: Comprehensive benchmark suite for video generative models
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., Wang, Y., Chen, X., Wang, L., Lin, D., Qiao, Y., and Liu, Z · 2024
Closest in time.
Vmc: Video motion customization using temporal attention adaption for text-to-video diffusion models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meiri, B., Samuel, D., Darshan, N., Chechik, G., Avidan, S., and Ben-Ari, R · 2023
Cited alongside, same era.
Localizing object-level shape variations with text-to-image diffusion models
Patashnik, O., Garibi, D., Azuri, I., Averbuch-Elor, H., and Cohen-Or, D · 2023
Cited alongside, same era.
Fatezero: Fusing attentions for zero-shot text-based video editing
Qi, C., Cun, X., Zhang, Y., Lei, C., Wang, X., Shan, Y., and Chen, Q · 2023
Cited alongside, same era.
Low-rank adaptation for fast text-to-image diffusion fine-tuning
Ryu, S · 2023
Cited alongside, same era.
Key-locked rank one editing for text-to-image personalization
Tewel, Y., Gal, R., Chechik, G., and Atzmon, Y · 2023
Cited alongside, same era.
Motioneditor: Editing video motion via content-aware diffusion
Tu, S., Dai, Q., Cheng, Z.-Q., Hu, H., Han, X., Wu, Z., and Jiang, Y.-G · 2023
Cited alongside, same era.
Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation
Wei, Y., Zhang, Y., Ji, Z., Bai, J., Zhang, L., and Zuo, W · 2023
Cited alongside, same era.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Wu, J. Z., Ge, Y., Wang, X., Lei, S. W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z · 2023
Cited alongside, same era.
Jeong, H., Park, G. Y., and Ye, J. C · 2024
Closest in time.
Videobooth: Diffusion-based video generation with image prompts
Jiang, Y., Wu, T., Yang, S., Si, C., Lin, D., Qiao, Y., Loy, C. C., and Liu, Z · 2024
Closest in time.
Video-p2p: Video editing with cross-attention control
Liu, S., Zhang, Y., Li, W., Lin, Z., and Jia, J · 2024
Closest in time.
Customize-a-video: One-shot motion customization of text-to-video diffusion models
Ren, Y., Zhou, Y., Yang, J., Shi, J., Liu, D., Liu, F., Kwon, M., and Shrivastava, A · 2024
Closest in time.
Edit-a-video: Single video editing with object-aware consistency
Shin, C., Kim, H., Lee, C. H., Lee, S.-g., and Yoon, S · 2024
Closest in time.
Training-free consistent text-to-image generation
Tewel, Y., Kaduri, O., Gal, R., Kasten, Y., Wolf, L., Chechik, G., and Atzmon, Y · 2024
Closest in time.
Motionbooth: Motion-aware customized text-to-video generation
Wu, J., Li, X., Zeng, Y., Zhang, J., Zhou, Q., Li, Y., Tong, Y., and Chen, K · 2024
Closest in time.
Space-time diffusion features for zero-shot text-driven motion transfer
Yatim, D., Fridman, R., Bar-Tal, O., Kasten, Y., and Dekel, T · 2024
Closest in time.
Jedi: Joint-image diffusion models for finetuning-free personalized text-to-image generation
Zeng, Y., Patel, V. M., Wang, H., Huang, X., Wang, T.-C., Liu, M.-Y., and Balaji, Y · 2024
Closest in time.
Motiondirector: Motion customization of text-to-video diffusion models
Zhao, R., Gu, Y., Wu, J. Z., Zhang, D. J., Liu, J.-W., Wu, W., Keppo, J., and Shou, M. Z · 2024
Closest in time.
Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise
Burgert, R., Xu, Y., Xian, W., Pilarski, O., Clausen, P., He, M., Ma, L., Deng, Y., Li, L., Mousavi, M., Ryoo, M., Debevec, P., and Yu, N · 2025
Closest in time.
Motionclone: Training-free motion cloning for controllable video generation
Ling, P., Bu, J., Zhang, P., Dong, X., Zang, Y., Wu, T., Chen, H., Wang, J., and Jin, Y · 2025
Closest in time.
Equivdm: Equivariant video diffusion models with temporally consistent noise, 2025
Liu, C. and Vahdat, A · 2025
Closest in time.
Wan: Open and advanced large-scale video generative models
Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.-W., Chen, D., Yu, F., Zhao, H., Yang, J., Zeng, J., Wang, J., Zhang, J., Zhou, J., Wang, J., Chen, J., Zhu, K., Zhao, K., Yan, K., Huang, L., Feng, M., Zhang, N., Li, P., Wu, P., Chu, R., Feng, R., Zhang, S., Sun, S., Fang, T., Wang, T., Gui, T., Weng, T., Shen, T., Lin, W., Wang, W., Wang, W., Zhou, W., Wang, W., Shen, W., Yu, W., Shi, X., Huang, X., Xu, X., Kou, Y., Lv, Y., Li, Y., Liu, Y., Wang, Y., Zhang, Y., Huang, Y., Li, Y., Wu, Y., Liu, Y., Pan, Y., Zheng, Y., Hong, Y., Shi, Y., Feng, Y., Jiang, Z., Han, Z., Wu, Z.-F., and Liu, Z · 2025
Closest in time.
Identity-preserving text-to-video generation by frequency decomposition
Yuan, S., Huang, J., He, X., Ge, Y., Shi, Y., Chen, L., Luo, J., and Yuan, L · 2025
Closest in time.
Magic mirror: Id-preserved video generation in video diffusion transformers
Zhang, Y., Liu, Y., Xia, B., Peng, B., Yan, Z., Lo, E., and Jia, J · 2025
Closest in time.