Fetching the paper…
Reading the bibliography…
Recent years have seen substantial progress in diffusion-based controllable video generation.
Some methods for classification and analysis of multivariate observations
MacQueen, J.; et al. 1967 · 1967
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2020 · 2011
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T.; Van Steenkiste, S.; Kurach, K.; Marinier, R.; Michalski, M.; and Gelly, S. 2018 · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P.; and Nichol, A. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Video diffusion models
Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; and Fleet, D. J. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Earlier work this paper cites.
Improving image generation with better captions
Betker, J.; Goh, G.; Jing, L.; Brooks, T.; Wang, J.; Li, L.; Ouyang, L.; Zhuang, J.; Lee, J.; Guo, Y.; et al. 2023 · 2023
Earlier work this paper cites.
Pix2video: Video editing using image diffusion
Ceylan, D.; Huang, C.-H. P.; and Mitra, N. J. 2023 · 2023
Cited alongside, same era.
Tracking anything with decoupled video segmentation
Cheng, H. K.; Oh, S. W.; Price, B.; Schwing, A.; and Lee, J.-Y. 2023 · 2023
Cited alongside, same era.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Hu, L.; Gao, X.; Zhang, P.; Sun, K.; Zhang, B.; and Bo, L. 2023 · 2023
Cited alongside, same era.
VideoBooth: Diffusion-based Video Generation with Image Prompts
Jiang, Y.; Wu, T.; Yang, S.; Si, C.; Lin, D.; Qiao, Y.; Loy, C. C.; and Liu, Z. 2023 · 2023
Cited alongside, same era.
Cotracker: It is better to track together
Karaev, N.; Rocco, I.; Graham, B.; Neverova, N.; Vedaldi, A.; and Rupprecht, C. 2023 · 2023
Magicanimate: Temporally consistent human image animation using diffusion model
Xu, Z.; Zhang, J.; Liew, J. H.; Yan, H.; Liu, J.-W.; Zhang, C.; Feng, J.; and Shou, M. Z. 2023 · 2023
Later among the works it cites.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Ye, H.; Zhang, J.; Liu, S.; Han, X.; and Yang, W. 2023 · 2023
Later among the works it cites.
Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory
Yin, S.; Wu, C.; Liang, J.; Shi, J.; Li, H.; Ming, G.; and Duan, N. 2023 · 2023
Later among the works it cites.
Make pixels dance: High-dynamic video generation
Zeng, Y.; Wei, G.; Zheng, J.; Zou, J.; Wei, Y.; Zhang, Y.; and Li, H. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dreampose: Fashion image-to-video synthesis via stable diffusion
Karras, J.; Holynski, A.; Wang, T.-C.; and Kemelmacher-Shlizerman, I. 2023 · 2023
Cited alongside, same era.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
Khachatryan, L.; Movsisyan, A.; Tadevosyan, V.; Henschel, R.; Wang, Z.; Navasardyan, S.; and Shi, H. 2023 · 2023
Cited alongside, same era.
TrailBlazer: Trajectory Control for Diffusion-Based Video Generation
Ma, W.-D. K.; Lewis, J.; and Kleijn, W. B. 2023 · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W.; and Xie, S. 2023 · 2023
Cited alongside, same era.
Fatezero: Fusing attentions for zero-shot text-based video editing
Qi, C.; Cun, X.; Zhang, Y.; Lei, C.; Wang, X.; Shan, Y.; and Chen, Q. 2023 · 2023
Cited alongside, same era.
MotionEditor: Editing Video Motion via Content-Aware Diffusion
Tu, S.; Dai, Q.; Cheng, Z.-Q.; Hu, H.; Han, X.; Wu, Z.; and Jiang, Y.-G. 2023 · 2023
Cited alongside, same era.
Dynamicrafter: Animating open-domain images with video diffusion priors
Xing, J.; Xia, M.; Zhang, Y.; Chen, H.; Wang, X.; Wong, T.-T.; and Shan, Y. 2023 · 2023
Cited alongside, same era.
Zhang, L.; Rao, A.; and Agrawala, M. 2023 · 2023
Later among the works it cites.
I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models
Zhang, S.; Wang, J.; Zhang, Y.; Zhao, K.; Yuan, H.; Qin, Z.; Wang, X.; Zhao, D.; and Zhou, J. 2023 · 2023
Later among the works it cites.
Video generation models as world simulators
Brooks, T.; Peebles, B.; Holmes, C.; DePue, W.; Guo, Y.; Jing, L.; Schnurr, D.; Taylor, J.; Luhman, T.; Luhman, E.; Ng, C.; Wang, R.; and Ramesh, A. 2024 · 2024
Closest in time.
Chen, J.; Ge, C.; Xie, E.; Wu, Y.; Yao, L.; Ren, X.; Wang, Z.; Luo, P.; Lu, H.; and Li, Z. 2024 · 2024
Closest in time.
Streamingt2v: Consistent, dynamic, and extendable long video generation from text
Henschel, R.; Khachatryan, L.; Hayrapetyan, D.; Poghosyan, H.; Tadevosyan, V.; Wang, Z.; Navasardyan, S.; and Shi, H. 2024 · 2024
Closest in time.
Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts
Ma, Y.; He, Y.; Wang, H.; Wang, A.; Qi, C.; Cai, C.; Li, X.; Li, Z.; Shum, H.-Y.; Liu, W.; et al. 2024 · 2024
Closest in time.
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Mou, C.; Wang, X.; Xie, L.; Wu, Y.; Zhang, J.; Qi, Z.; and Shan, Y. 2024 · 2024
Closest in time.
DragAnything: Motion Control for Anything using Entity Representation
Wu, W.; Li, Z.; Gu, Y.; Zhao, R.; He, Y.; Zhang, D. J.; Shou, M. Z.; Li, Y.; Gao, T.; and Zhang, D. 2024 · 2024
Closest in time.