Fetching the paper…
Reading the bibliography…
Despite impressive advancements in diffusion-based video editing models in altering video attributes, there has been limited exploration into modifying motion information while preserving the original protagonist's appearance and background.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
Image quality metrics: Psnr vs. ssim
A. Hore and D. Ziou · 2010
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
Conditional gan with discriminative filter generation for text-to-video synthesis
Y. Balaji, M. R. Min, B. Bai, R. Chellappa, and H. P. Graf · 2019
Earlier work this paper cites.
Liquid warping gan: A unified framework for human motion imitation, appearance transfer and novel view synthesis
W. Liu, Z. Piao, J. Min, W. Luo, L. Ma, and S. Gao · 2019
Earlier work this paper cites.
First order motion model for image animation
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Improved techniques for training score-based generative models
Y. Song and S. Ermon · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Earlier work this paper cites.
Few-shot human motion transfer by personalized geometry and texture modeling
Z. Huang, X. Han, J. Xu, and T. Zhang · 2021
Earlier work this paper cites.
Learning high fidelity depths of dressed humans by watching social media dance videos
Y. Jafarian and H. S. Park · 2021
Earlier work this paper cites.
Sdedit: Guided image synthesis and editing with stochastic differential equations
C. Meng, Y. He, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
A. Q. Nichol and P. Dhariwal · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Motion representations for articulated animation
A. Siarohin, O. J. Woodford, J. Ren, M. Chai, and S. Tulyakov · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2021
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2021
Cited alongside, same era.
Text2live: Text-driven layered image and video editing
O. Bar-Tal, D. Ofri-Amar, R. Fridman, Y. Kasten, and T. Dekel · 2022
Cited alongside, same era.
Cascaded diffusion models for high fidelity image generation
J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans · 2022
Cited alongside, same era.
Video diffusion models
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Pix2video: Video editing using image diffusion
D. Ceylan, C.-H. P. Huang, and N. J. Mitra · 2023
Cited alongside, same era.
Fatezero: Fusing attentions for zero-shot text-based video editing
C. Qi, X. Cun, Y. Zhang, C. Lei, X. Wang, Y. Shan, and Q. Chen · 2023
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, et al · 2023
Later among the works it cites.
Zero-shot video editing using off-the-shelf image diffusion models
W. Wang, K. Xie, Z. Liu, H. Chen, Y. Cao, X. Wang, and C. Shen · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
J. Z. Wu, Y. Ge, X. Wang, S. W. Lei, Y. Gu, Y. Shi, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou · 2023
Later among the works it cites.
Open-vocabulary panoptic segmentation with text-to-image diffusion models
J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stablevideo: Text-driven consistency-aware diffusion video editing
W. Chai, X. Guo, G. Wang, and Y. Lu · 2023
Cited alongside, same era.
Flatten: optical flow-guided attention for consistent text-to-video editing
Y. Cong, M. Xu, C. Simon, S. Chen, J. Ren, Y. Xie, J.-M. Perez-Rua, B. Rosenhahn, T. Xiang, and S. He · 2023
Cited alongside, same era.
Videdit: Zero-shot and spatially aware text-driven video editing
P. Couairon, C. Rambour, J.-E. Haugeard, and N. Thome · 2023
Cited alongside, same era.
Structure and content-guided video synthesis with diffusion models
P. Esser, J. Chiu, P. Atighehchian, J. Granskog, and A. Germanidis · 2023
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang · 2023
Cited alongside, same era.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
L. Khachatryan, A. Movsisyan, V. Tadevosyan, R. Henschel, Z. Wang, S. Navasardyan, and H. Shi · 2023
Cited alongside, same era.
Effective whole-body pose estimation with two-stages distillation
Z. Yang, A. Zeng, C. Yuan, and Y. Li · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
Later among the works it cites.
Ccedit: Creative and controllable video editing via diffusion models
R. Feng, W. Weng, Y. Wang, Y. Yuan, J. Bao, C. Luo, Z. Chen, and B. Guo · 2024
Closest in time.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Y. Guo, C. Yang, A. Rao, Y. Wang, Y. Qiao, D. Lin, and B. Dai · 2024
Closest in time.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
L. Hu, X. Gao, P. Zhang, K. Sun, B. Zhang, and L. Bo · 2024
Closest in time.
Follow your pose: Pose-guided text-to-video generation using pose-free videos
Y. Ma, Y. He, X. Cun, X. Wang, Y. Shan, X. Li, and Q. Chen · 2024
Closest in time.
Dragondiffusion: Enabling drag-style manipulation on diffusion models
C. Mou, X. Wang, J. Song, Y. Shan, and J. Zhang · 2024
Closest in time.
Motioneditor: Editing video motion via content-aware diffusion
S. Tu, Q. Dai, Z.-Q. Cheng, H. Hu, X. Han, Z. Wu, and Y.-G. Jiang · 2024
Closest in time.
Disco: Disentangled control for referring human dance generation in real world
T. Wang, L. Li, K. Lin, C.-C. Lin, Z. Yang, H. Zhang, Z. Liu, and L. Wang · 2024
Closest in time.
Magicanimate: Temporally consistent human image animation using diffusion model
Z. Xu, J. Zhang, J. H. Liew, H. Yan, J.-W. Liu, C. Zhang, J. Feng, and M. Z. Shou · 2024
Closest in time.
Controlvideo: Training-free controllable text-to-video generation
Y. Zhang, Y. Wei, D. Jiang, X. Zhang, W. Zuo, and Q. Tian · 2024
Closest in time.
Champ: Controllable and consistent human image animation with 3d parametric guidance
S. Zhu, J. L. Chen, Z. Dai, Y. Xu, X. Cao, Y. Yao, H. Zhu, and S. Zhu · 2024
Closest in time.