Fetching the paper…
Reading the bibliography…
Recent diffusion-based human image animation techniques have demonstrated impressive success in synthesizing videos that faithfully follow a given reference identity and a sequence of desired movement poses.
Self-supervised learning for semi-supervised temporal action proposal
Wang X, Zhang S, Qing Z, et al · 1914
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Wang Z, Bovik A C, Sheikh H R, et al · 2004
Earlier work this paper cites.
Image quality metrics: Psnr vs. ssim
Hore A and Ziou D · 2010
Earlier work this paper cites.
Generative adversarial nets
Goodfellow I, Pouget-Abadie J, Mirza M, et al · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger O, Fischer P, and Brox T · 2015
Earlier work this paper cites.
Temporal convolutional networks for action segmentation and detection
Lea C, Flynn M D, Vidal R, et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov I and Hutter F · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Qiu Z, Yao T, and Mei T · 2017
Earlier work this paper cites.
Video generation from text
Li Y, Min M, Shen D, et al · 2018
Earlier work this paper cites.
MocoGAN: Decomposing motion and content for video generation
Tulyakov S, Liu M Y, Yang X, et al · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner T, Van Steenkiste S, Kurach K, et al · 2018
Earlier work this paper cites.
Pose guided human video generation
Yang C, Wang Z, Zhu X, et al · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang R, Isola P, Efros A A, et al · 2018
Earlier work this paper cites.
Dense intrinsic appearance flow for human pose transfer
Li Y, Huang C, and Loy C C · 2019
Earlier work this paper cites.
First order motion model for image animation
Siarohin A, Lathuilière S, Tulyakov S, et al · 2019
Earlier work this paper cites.
Dwnet: Dense warp-based network for pose-guided human video generation
Zablotskaia P, Siarohin A, Zhao B, et al · 2019
Earlier work this paper cites.
Dwnet: Dense warp-based network for pose-guided human video generation
Zablotskaia P, Siarohin A, Zhao B, et al · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho J, Jain A, and Abbeel P · 2020
Earlier work this paper cites.
G3an: Disentangling appearance and motion for video generation
Wang Y, Bilinski P, Bremond F, et al · 2020
Earlier work this paper cites.
Vivit: A video vision transformer
Arnab A, Dehghani M, Heigold G, et al · 2021
Earlier work this paper cites.
Is space-time attention all you need for video understanding?
Bertasius G, Wang H, and Torresani L · 2021
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
Gu A, Goel K, and Ré C · 2021
Earlier work this paper cites.
Learning high fidelity depths of dressed humans by watching social media dance videos
Jafarian Y and Park H S · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford A, Kim J W, Hallacy C, et al · 2021
Earlier work this paper cites.
Motion representations for articulated animation
Siarohin A, Woodford O J, Ren J, et al · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
Song J, Meng C, and Ermon S · 2021
Earlier work this paper cites.
Oadtr: Online action detection with transformers
Wang X, Zhang S, Qing Z, et al · 2021
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
Ho J, Chan W, Saharia C, et al · 2022
Earlier work this paper cites.
Text2human: Text-driven controllable human image generation
Jiang Y, Yang S, Qiu H, et al · 2022
Earlier work this paper cites.
Transcrowd: weakly-supervised crowd counting with transformers
Liang D, Chen X, Xu W, et al · 2022
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol A Q, Dhariwal P, Ramesh A, et al · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Ramesh A, Dhariwal P, Nichol A, et al · 2022
Earlier work this paper cites.
Neural texture extraction and distribution for controllable person image synthesis
Ren Y, Fan X, Li G, et al · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach R, Blattmann A, Lorenz D, et al · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia C, Chan W, Saxena S, et al · 2022
Cited alongside, same era.
Exploring dual-task correlation for pose guided person image generation
Zhang P, Yang L, Lai J H, et al · 2022
Cited alongside, same era.
Exploring dual-task correlation for pose guided person image generation
Zhang P, Yang L, Lai J H, et al · 2022
Cited alongside, same era.
Thin-plate spline motion model for image animation
Zhao J and Zhang H · 2022
Cited alongside, same era.
Videocomposer: Compositional video synthesis with motion controllability
Wang X, Yuan H, Zhang S, et al · 2023
Later among the works it cites.
Few-shot action recognition with captioning foundation models
Wang X, Zhang S, Yuan H, et al · 2023
Later among the works it cites.
Videolcm: Video latent consistency model
Wang X, Zhang S, Zhang H, et al · 2023
Later among the works it cites.
Dreamvideo: Composing your dream videos with customized subject and motion
Wei Y, Zhang S, Qing Z, et al · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Wu J Z, Ge Y, Wang X, et al · 2023
Later among the works it cites.
Make-your-video: Customized video generation using textual and structural guidance
Xing J, Xia M, Liu Y, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Magicvideo: Efficient video generation with latent diffusion models
Zhou D, Wang W, Yan H, et al · 2022
Cited alongside, same era.
Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation
An J, Zhang S, Yang H, et al · 2023
Cited alongside, same era.
Person image synthesis via denoising diffusion model
Bhunia A K, Khan S, Cholakkal H, et al · 2023
Cited alongside, same era.
Align your latents: High-resolution video synthesis with latent diffusion models
Blattmann A, Rombach R, Ling H, et al · 2023
Cited alongside, same era.
Pix2video: Video editing using image diffusion
Ceylan D, Huang C H P, and Mitra N J · 2023
Cited alongside, same era.
Stablevideo: Text-driven consistency-aware diffusion video editing
Chai W, Guo X, Wang G, et al · 2023
Cited alongside, same era.
Later among the works it cites.
Simda: Simple diffusion adapter for efficient video generation
Xing Z, Dai Q, Hu H, et al · 2023
Later among the works it cites.
Magicanimate: Temporally consistent human image animation using diffusion model
Xu Z, Zhang J, Liew J H, et al · 2023
Later among the works it cites.
Effective whole-body pose estimation with two-stages distillation
Yang Z, Zeng A, Yuan C, et al · 2023
Later among the works it cites.
Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory
Yin S, Wu C, Liang J, et al · 2023
Later among the works it cites.
Bidirectionally deformable motion modulation for video-based human pose transfer
Yu W Y, Po L M, Cheung R C, et al · 2023
Later among the works it cites.
Instructvideo: Instructing video diffusion models with human feedback
Yuan H, Zhang S, Wang X, et al · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang L, Rao A, and Agrawala M · 2023
Later among the works it cites.
I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models
Zhang S, Wang J, Zhang Y, et al · 2023
Later among the works it cites.
Controlvideo: Training-free controllable text-to-video generation
Zhang Y, Wei Y, Jiang D, et al · 2023
Later among the works it cites.
Controlvideo: Adding conditional control for one shot text-to-video editing
Zhao M, Wang R, Bao F, et al · 2023
Later among the works it cites.
Video mamba suite: State space model as a versatile alternative for video understanding
Chen G, Huang Y, Xu J, et al · 2024
Closest in time.
Videomamba: State space model for efficient video understanding
Li K, Li X, Wang Y, et al · 2024
Closest in time.
Monkey: Image resolution and text label are important things for large multi-modal models
Li Z, Yang B, Liu Q, et al · 2024
Closest in time.
Vmamba: Visual state space model
Liu Y, Tian Y, Zhao Y, et al · 2024
Closest in time.
Textmonkey: An ocr-free large multimodal model for understanding document
Liu Y, Yang B, Liu Q, et al · 2024
Closest in time.
Follow your pose: Pose-guided text-to-video generation using pose-free videos
Ma Y, He Y, Cun X, et al · 2024
Closest in time.
Cpt: A pre-trained unbalanced transformer for both chinese language understanding and generation
Shao Y, Geng Z, Liu Y, et al · 2024
Closest in time.
Disco: Disentangled control for referring human dance generation in real world
Wang T, Li L, Lin K, et al · 2024
Closest in time.
A recipe for scaling up text-to-video generation with text-free videos
Wang X, Zhang S, Yuan H, et al · 2024
Closest in time.
Do you guys want to dance: Zero-shot compositional human dance generation with multiple persons
Xu Z, Wei K, Yang X, et al · 2024
Closest in time.
Plainmamba: Improving non-hierarchical mamba in visual recognition
Yang C, Chen Z, Espinosa M, et al · 2024
Closest in time.
Poseanimate: Zero-shot high fidelity pose controllable character animation
Zhu B, Wang F, Lu T, et al · 2024
Closest in time.
Vision mamba: Efficient visual representation learning with bidirectional state space model
Zhu L, Liao B, Zhang Q, et al · 2024
Closest in time.
Champ: Controllable and consistent human image animation with 3d parametric guidance
Zhu S, Chen J L, Dai Z, et al · 2024
Closest in time.
Rv-gan: Recurrent gan for unconditional video generation
Gupta S, Keshari A, and Das S · 2033
Closest in time.