Fetching the paper…
Reading the bibliography…
Recent advances in video diffusion models have enabled realistic and controllable human image animation with temporal coherence.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
Image quality metrics: Psnr vs. ssim
A. Hore and D. Ziou · 2010
Earlier work this paper cites.
High-quality passive facial performance capture using anchor frames
T. Beeler, F. Hahn, D. Bradley, B. Bickel, P. Beardsley, C. Gotsman, R. W. Sumner, and M. Gross · 2011
Earlier work this paper cites.
Multiview face capture using polarized spherical gradient illumination
A. Ghosh, G. Fyffe, B. Tunwattanapong, J. Busch, X. Yu, and P. Debevec · 2011
Earlier work this paper cites.
High-quality streamable free-viewpoint video
A. Collet, M. Chuang, P. Sweeney, D. Gillett, D. Evseev, D. Calabrese, H. Hoppe, A. Kirk, and S. Sullivan · 2015
Earlier work this paper cites.
SMPL: A skinned multi-person linear model
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black · 2015
Earlier work this paper cites.
Realtime multi-person 2d pose estimation using part affinity fields
Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Densepose: Dense human pose estimation in the wild
R. A. Güler, N. Neverova, and I. Kokkinos · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
Conditional gan with discriminative filter generation for text-to-video synthesis
Y. Balaji, M. R. Min, B. Bai, R. Chellappa, and H. P. Graf · 2019
Earlier work this paper cites.
Everybody dance now
C. Chan, S. Ginosar, T. Zhou, and A. A. Efros · 2019
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition
J. Deng, J. Guo, N. Xue, and S. Zafeiriou · 2019
Earlier work this paper cites.
Towards multi-pose guided virtual try-on network
H. Dong, X. Liang, X. Shen, B. Wang, H. Lai, J. Zhu, Z. Hu, and J. Yin · 2019
Earlier work this paper cites.
3d guided fine-grained face manipulation
Z. Geng, C. Cao, and S. Tulyakov · 2019
Earlier work this paper cites.
The relightables: Volumetric performance capture of humans with realistic relighting
K. Guo, P. Lincoln, P. Davidson, J. Busch, X. Yu, M. Whalen, G. Harvey, S. Orts-Escolano, R. Pandey, J. Dourgarian, et al · 2019
Cited alongside, same era.
Expressive body capture: 3d hands, face, and body from a single image
G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. A. Osman, D. Tzionas, and M. J. Black · 2019
Cited alongside, same era.
First order motion model for image animation
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe · 2019
Cited alongside, same era.
Learning high fidelity depths of dressed humans by watching social media dance videos
Y. Jafarian and H. S. Park · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
On the importance of noise scheduling for diffusion models
T. Chen · 2023
Later among the works it cites.
Dreamoving: A human dance video generation framework based on diffusion models
M. Feng, J. Liu, K. Yu, Y. Yao, Z. Hui, X. Guo, X. Lin, H. Xue, C. Shi, X. Li, et al · 2023
Later among the works it cites.
EVA3d: Compositional 3d human generation from 2d image collections
F. Hong, Z. Chen, Y. LAN, L. Pan, and Z. Liu · 2023
Later among the works it cites.
Rtmpose: Real-time multi-person pose estimation based on mmpose
T. Jiang, P. Lu, L. Zhang, N. Ma, R. Han, C. Lyu, Y. Li, and K. Chen · 2023
Later among the works it cites.
Dreampose: Fashion image-to-video synthesis via stable diffusion
J. Karras, A. Holynski, T.-C. Wang, and I. Kemelmacher-Shlizerman · 2023
Later among the works it cites.
Dinar: Diffusion inpainting of neural textures for one-shot human avatars
D. Svitov, D. Gudkov, R. Bashirov, and V. Lempitsky · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Siarohin, O. Woodford, J. Ren, M. Chai, and S. Tulyakov · 2021
Cited alongside, same era.
3d pose transfer with correspondence learning and mesh refinement
C. Song, J. Wei, R. Li, F. Liu, and G. Lin · 2021
Cited alongside, same era.
One-shot free-view neural talking-head synthesis for video conferencing
T.-C. Wang, A. Mallya, and M.-Y. Liu · 2021
Cited alongside, same era.
Modeling clothing as a separate layer for an animatable human avatar
D. Xiang, F. Prada, T. Bagautdinov, W. Xu, Y. Dong, H. Wen, J. Hodgins, and C. Wu · 2021
Cited alongside, same era.
Pose-guided human animation from a single image in the wild
J. S. Yoon, L. Liu, V. Golyanik, K. Sarkar, H. S. Park, and C. Theobalt · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Progressive distillation for fast sampling of diffusion models
T. Salimans and J. Ho · 2022
Cited alongside, same era.
Later among the works it cites.
Disco: Disentangled control for referring human dance generation in real world
T. Wang, L. Li, K. Lin, C.-C. Lin, Z. Yang, H. Zhang, Z. Liu, and L. Wang · 2023
Later among the works it cites.
Xagen: 3d expressive human avatars generation
Z. Xu, J. Zhang, J. H. Liew, J. Feng, and M. Z. Shou · 2023
Later among the works it cites.
Effective whole-body pose estimation with two-stages distillation
Z. Yang, A. Zeng, C. Yuan, and Y. Li · 2023
Later among the works it cites.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
H. Ye, J. Zhang, S. Liu, X. Han, and W. Yang · 2023
Later among the works it cites.
Bidirectionally deformable motion modulation for video-based human pose transfer
W.-Y. Yu, L.-M. Po, R. C. Cheung, Y. Zhao, Y. Xue, and K. Li · 2023
Later among the works it cites.
Wear-any-way: Manipulable virtual try-on via sparse correspondence alignment
M. Chen, X. Chen, Z. Zhai, C. Ju, X. Hong, J. Lan, and S. Xiao · 2024
Closest in time.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Y. Guo, C. Yang, A. Rao, Y. Wang, Y. Qiao, D. Lin, and B. Dai · 2024
Closest in time.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
L. Hu, X. Gao, P. Zhang, K. Sun, B. Zhang, and L. Bo · 2024
Closest in time.
Common diffusion noise schedules and sample steps are flawed
S. Lin, B. Liu, J. Li, and X. Yang · 2024
Closest in time.
Champ: Controllable and consistent human image animation with 3d parametric guidance
S. Zhu, J. L. Chen, Z. Dai, Y. Xu, X. Cao, Y. Yao, H. Zhu, and S. Zhu · 2024
Closest in time.