Fetching the paper…
Reading the bibliography…
Diffusion models have achieved remarkable progress on image-to-video (I2V) generation, while their noise-to-data generation process is inherently mismatched with this task, which may lead to suboptimal synthesis quality.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K · 2012
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
Temporal generative adversarial nets with singular value clipping
Saito, M., Matsumoto, E., and Saito, S · 2017
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Bain, M., Nagrani, A., Varol, G., and Zisserman, A · 2021
Earlier work this paper cites.
Diffusion schrödinger bridge with applications to score-based generative modeling
De Bortoli, V., Thornton, J., Heng, J., and Doucet, A · 2021
Earlier work this paper cites.
Variational diffusion models
Kingma, D., Salimans, T., Poole, B., and Ho, J · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Score-based generative modeling in latent space
Vahdat, A., Kreis, K., and Kautz, J · 2021
Earlier work this paper cites.
Godiva: Generating open-domain videos from natural descriptions
Wu, C., Huang, L., Zhang, Q., Li, B., Ji, L., Yang, F., Sapiro, G., and Duan, N · 2021
Earlier work this paper cites.
Latent video diffusion models for high-fidelity long video generation
He, Y., Yang, T., Zhang, Y., Shan, Y., and Chen, Q · 2022
Earlier work this paper cites.
Make it move: controllable image-to-video generation with text descriptions
Hu, Y., Luo, C., and Chen, Z · 2022
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A. Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., Mcgrew, B., Sutskever, I., and Chen, M · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents, 2022
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Skorokhodov, I., Tulyakov, S., and Elhoseiny, M · 2022
Cited alongside, same era.
Stochastic interpolants: A unifying framework for flows and diffusions
Albergo, M. S., Boffi, N. M., and Vanden-Eijnden, E · 2023
Cited alongside, same era.
All are worth words: A vit backbone for diffusion models
Bao, F., Nie, S., Xue, K., Cao, Y., Li, C., Su, H., and Zhu, J · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al · 2023
Cited alongside, same era.
Boosting latent diffusion with flow matching
Fischer, J. S., Gui, M., Ma, P., Stracke, N., Baumann, S. A., and Ommer, B · 2023
Cited alongside, same era.
On the content bias in fréchet video distance
Ge, S., Mahapatra, A., Parmar, G., Zhu, J.-Y., and Huang, J.-B · 2024
Closest in time.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Liang, Z., Wang, Y., Qiao, Y., Agrawala, M., Lin, D., and Dai, B · 2024
Closest in time.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Hu, L · 2024
Closest in time.
Video interpolation with diffusion models
Jain, S., Watson, D., Tabellion, E., Poole, B., Kontkanen, J., et al · 2024
Closest in time.
Generative image dynamics
Li, Z., Tucker, R., Snavely, N., and Holynski, A · 2024
Closest in time.
Common diffusion noise schedules and sample steps are flawed
Lin, S., Liu, B., Li, J., and Yang, X · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dreamsim: Learning new dimensions of human visual similarity using synthetic data
Fu, S., Tamir, N., Sundaram, S., Chai, L., Zhang, R., Dekel, T., and Isola, P · 2023
Cited alongside, same era.
Preserve your own correlation: A noise prior for video diffusion models
Ge, S., Nah, S., Liu, G., Poon, T., Tao, A., Catanzaro, B., Jacobs, D., Huang, J.-B., Liu, M.-Y., and Balaji, Y · 2023
Cited alongside, same era.
I 2 sb: Image-to-image schrödinger bridge
Liu, G.-H., Vahdat, A., Huang, D.-A., Theodorou, E. A., Nie, W., and Anandkumar, A · 2023
Cited alongside, same era.
Conditional image-to-video generation with latent flow diffusion models
Ni, H., Shi, C., Li, K., Huang, S. X., and Min, M. R · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Non-denoising forward-time diffusions
Peluchetti, S · 2023
Cited alongside, same era.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Wu, J. Z., Ge, Y., Wang, X., Lei, S. W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z · 2023
Cited alongside, same era.
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling
Shi, X., Huang, Z., Wang, F.-Y., Bian, W., Li, D., Zhang, Y., Zhang, M., Cheung, K. C., See, S., Qin, H., et al · 2024
Closest in time.
Simulation-free schrödinger bridges via score and flow matching
Tong, A., Malkin, N., Fatras, K., Atanackovic, L., Zhang, Y., Huguet, G., Wolf, G., and Bengio, Y · 2024
Closest in time.
Videocomposer: Compositional video synthesis with motion controllability
Wang, X., Yuan, H., Zhang, S., Chen, D., Wang, J., Zhang, Y., Shen, Y., Zhao, D., and Zhou, J · 2024
Closest in time.
Freeinit: Bridging initialization gap in video diffusion models
Wu, T., Si, C., Jiang, Y., Huang, Z., and Liu, Z · 2024
Closest in time.
Dynamicrafter: Animating open-domain images with video diffusion priors
Xing, J., Xia, M., Zhang, Y., Chen, H., Yu, W., Liu, H., Liu, G., Wang, X., Shan, Y., and Wong, T.-T · 2024
Closest in time.
Efficient video diffusion models via content-frame motion-latent decomposition
Yu, S., Nie, W., Huang, D. A., Li, B., Shin, J., and Anandkumar, A · 2024
Closest in time.
Diffusion bridge implicit models
Zheng, K., He, G., Chen, J., Bao, F., and Zhu, J · 2024
Closest in time.
Sparsectrl: Adding sparse controls to text-to-video diffusion models
Guo, Y., Yang, C., Rao, A., Agrawala, M., Lin, D., and Dai, B · 2025
Closest in time.
Bridge-sr: Schrödinger bridge for efficient sr
Li, C., Chen, Z., Bao, F., and Zhu, J · 2025
Closest in time.
Lavie: High-quality video generation with cascaded latent diffusion models
Wang, Y., Chen, X., Ma, X., Zhou, S., Huang, Z., Wang, Y., Yang, C., He, Y., Yu, J., Yang, P., et al · 2025
Closest in time.