Fetching the paper…
Reading the bibliography…
The evolution of diffusion models has greatly impacted video generation and understanding.
Motion estimation using a complex-valued wavelet transform
Magarey, J.; and Kingsbury, N. 1998 · 1998
Earlier work this paper cites.
Highly scalable video compression using a lifting-based 3D wavelet transform with deformable mesh motion compensation
Secker, A.; and Taubman, D. 2002 · 2002
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Tweedie’s formula and selection bias
Efron, B. 2011 · 2011
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; and Ganguli, S. 2015 · 2015
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
Pont-Tuset, J.; Perazzi, F.; Caelles, S.; Arbeláez, P.; Sorkine-Hornung, A.; and Van Gool, L. 2017 · 2017
Earlier work this paper cites.
Deep convolutional framelets: A general deep learning framework for inverse problems
Ye, J. C.; Han, Y.; and Cha, E. 2018 · 2018
Earlier work this paper cites.
Photorealistic style transfer via wavelet transforms
Yoo, J.; Uh, Y.; Chun, S.; Kang, B.; and Ha, J.-W. 2019 · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Bain, M.; Nagrani, A.; Varol, G.; and Zisserman, A. 2021 · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J.; Holtzman, A.; Forbes, M.; Bras, R. L.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Delving into the frequency: Temporally consistent human motion transfer in the fourier space
Yang, G.; Liu, W.; Liu, X.; Gu, X.; Cao, J.; and Li, J. 2022 · 2022
Cited alongside, same era.
Pix2video: Video editing using image diffusion
Ceylan, D.; Huang, C.-H. P.; and Mitra, N. J. 2023 · 2023
Cited alongside, same era.
Hu, Z.; and Xu, D. 2023 · 2023
Cited alongside, same era.
VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion Models
Jeong, H.; Park, G. Y.; and Ye, J. C. 2023 · 2023
Cited alongside, same era.
Dreamvideo: Composing your dream videos with customized subject and motion
Wei, Y.; Zhang, S.; Qing, Z.; Yuan, H.; Liu, Z.; Liu, Y.; Zhang, Y.; Zhou, J.; and Shan, H. 2023 · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Wu, J. Z.; Ge, Y.; Wang, X.; Lei, S. W.; Gu, Y.; Shi, Y.; Hsu, W.; Shan, Y.; Qie, X.; and Shou, M. Z. 2023 · 2023
Later among the works it cites.
Space-Time Diffusion Features for Zero-Shot Text-Driven Motion Transfer
Yatim, D.; Fridman, R.; Tal, O. B.; Kasten, Y.; and Dekel, T. 2023 · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang, L.; Rao, A.; and Agrawala, M. 2023 · 2023
Later among the works it cites.
Motiondirector: Motion customization of text-to-video diffusion models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ground-a-video: Zero-shot grounded video editing using text-to-image diffusion models
Jeong, H.; and Ye, J. C. 2023 · 2023
Cited alongside, same era.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
Khachatryan, L.; Movsisyan, A.; Tadevosyan, V.; Henschel, R.; Wang, Z.; Navasardyan, S.; and Shi, H. 2023 · 2023
Cited alongside, same era.
Gligen: Open-set grounded text-to-image generation
Li, Y.; Liu, H.; Wu, Q.; Mu, F.; Yang, J.; Gao, J.; Li, C.; and Lee, Y. J. 2023 · 2023
Cited alongside, same era.
Fatezero: Fusing attentions for zero-shot text-based video editing
Qi, C.; Cun, X.; Zhang, Y.; Lei, C.; Wang, X.; Shan, Y.; and Chen, Q. 2023 · 2023
Cited alongside, same era.
Zeroscope
Sterling, S. 2023 · 2023
Cited alongside, same era.
Emergent correspondence from image diffusion
Tang, L.; Jia, M.; Wang, Q.; Phoo, C. P.; and Hariharan, B. 2023 · 2023
Cited alongside, same era.
Videocrafter1: Open diffusion models for high-quality video generation
Chen, H.; Xia, M.; He, Y.; Zhang, Y.; Cun, X.; Yang, S.; Xing, J.; Liu, Y.; Chen, Q.; Wang, X.; et al. 2023a
Cited in the paper.
Zhao, R.; Gu, Y.; Wu, J. Z.; Zhang, D. J.; Liu, J.; Wu, W.; Keppo, J.; and Shou, M. Z. 2023b · 2023
Later among the works it cites.
UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing
Bai, J.; He, T.; Wang, Y.; Guo, J.; Hu, H.; Liu, Z.; and Bian, J. 2024 · 2024
Closest in time.
DreamMotion: Space-Time Self-Similarity Score Distillation for Zero-Shot Video Editing
Jeong, H.; Chang, J.; Park, G. Y.; and Ye, J. C. 2024 · 2024
Closest in time.
Hybrid Video Diffusion Models with 2D Triplane and 3D Wavelet Representation
Kim, K.; Lee, H.; Park, J.; Kim, S.; Lee, K.; Kim, S.; and Yoo, J. 2024 · 2024
Closest in time.
Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models
Ren, Y.; Zhou, Y.; Yang, J.; Shi, J.; Liu, D.; Liu, F.; Kwon, M.; and Shrivastava, A. 2024 · 2024
Closest in time.
A Unified Framework for U-Net Design and Analysis
Williams, C.; Falck, F.; Deligiannidis, G.; Holmes, C. C.; Doucet, A.; and Syed, S. 2024 · 2024
Closest in time.
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
Yang, S.; Hou, L.; Huang, H.; Ma, C.; Wan, P.; Zhang, D.; Chen, X.; and Liao, J. 2024 · 2024
Closest in time.