Fetching the paper…
Reading the bibliography…
Human video generation is a dynamic and rapidly evolving task that aims to synthesize 2D human body video sequences with generative models given control conditions such as text, audio, and pose.
R. O. Cornett, “Cued speech,” American annals of the deaf , pp. 3–13, 1967
1967
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
L. Gorelick, M. Blank, E. Shechtman, M. Irani, and R. Basri, “Actions as space-time shapes,” IEEE transactions on pattern analysis and machine intelligence , vol. 29, no. 12, pp. 2247–2253, 2007
2007
Earlier work this paper cites.
L. Gorelick, M. Blank, E. Shechtman, M. Irani, and R. Basri, “Actions as space-time shapes,” IEEE transactions on pattern analysis and machine intelligence , vol. 29, no. 12, pp. 2247–2253, 2007
2007
Earlier work this paper cites.
A. Hore and D. Ziou, “Image quality metrics: Psnr vs. ssim,” in ICPR , 2010
2010
Earlier work this paper cites.
P. Lucey, J. F. Cohn, T. Kanade, J. Saragih, Z. Ambadar, and I. Matthews, “The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression,” in CVPR workshops , 2010
2010
Earlier work this paper cites.
Y. Yang and D. Ramanan, “Articulated human detection with flexible mixtures of parts,” IEEE transactions on pattern analysis and machine intelligence , vol. 35, no. 12, pp. 2878–2890, 2012
2012
Earlier work this paper cites.
K. Soomro, A. R. Zamir, and M. Shah, “Ucf101: A dataset of 101 human actions classes from videos in the wild,” arXiv , 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu, “Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments,” IEEE transactions on pattern analysis and machine intelligence , vol. 36, no. 7, pp. 1325–1339, 2013
2013
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” NeurIPS , 2014
2014
Earlier work this paper cites.
M. Mirza and S. Osindero, “Conditional generative adversarial nets,” arXiv , 2014
2014
Earlier work this paper cites.
N. Sadoughi, Y. Liu, and C. Busso, “Msp-avatar corpus: Motion capture recordings to study the role of discourse functions in the design of intelligent virtual agents,” in FG , 2015
2015
Earlier work this paper cites.
A. Kendall, M. Grimes, and R. Cipolla, “Posenet: A convolutional network for real-time 6-dof camera relocalization,” in ICCV , 2015
2015
Earlier work this paper cites.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,” NeurIPS , 2016
2016
Earlier work this paper cites.
A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+ d: A large scale dataset for 3d human activity analysis,” in CVPR , 2016
2016
Earlier work this paper cites.
Z. Liu, P. Luo, S. Qiu, X. Wang, and X. Tang, “Deepfashion: Powering robust clothes recognition and retrieval with rich annotations,” in CVPR , 2016
2016
Earlier work this paper cites.
A. Newell, K. Yang, and J. Deng, “Stacked hourglass networks for human pose estimation,” arXiv , 2016
2016
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” NeurIPS , 2017
2017
Earlier work this paper cites.
L. Chen, S. Srivastava, Z. Duan, and C. Xu, “Deep cross-modal audio-visual generation,” in ACM MM Workshops , 2017
2017
Earlier work this paper cites.
Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh, “Realtime multi-person 2d pose estimation using part affinity fields,” in CVPR , 2017
2017
Earlier work this paper cites.
E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox, “Flownet 2.0: Evolution of optical flow estimation with deep networks,” in CVPR , 2017
2017
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” NeurIPS , 2017
2017
Earlier work this paper cites.
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in CVPR , 2017
2017
Earlier work this paper cites.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in CVPR , 2018
2018
Earlier work this paper cites.
T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly, “Towards accurate generative models of video: A new metric & challenges,” arXiv , 2018
2018
Earlier work this paper cites.
S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz, “Mocogan: Decomposing motion and content for video generation,” in CVPR , 2018
2018
Earlier work this paper cites.
B. Li, X. Liu, K. Dinesh, Z. Duan, and G. Sharma, “Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,” IEEE Transactions on Multimedia , vol. 21, no. 2, pp. 522–535, 2018
2018
Earlier work this paper cites.
H. R. V. Joze and O. Koller, “Ms-asl: A large-scale data set and benchmark for understanding american sign language,” arXiv , 2018
2018
Earlier work this paper cites.
N. C. Camgoz, S. Hadfield, O. Koller, H. Ney, and R. Bowden, “Neural sign language translation,” in CVPR , 2018
2018
Earlier work this paper cites.
R. Mahjourian, M. Wicke, and A. Angelova, “Unsupervised learning of depth and ego-motion from monocular video using 3d geometric constraints,” in CVPR , 2018
2018
Earlier work this paper cites.
R. A. Güler, N. Neverova, and I. Kokkinos, “Densepose: Dense human pose estimation in the wild,” in CVPR , 2018
2018
Earlier work this paper cites.
C. Yang, Z. Wang, X. Zhu, C. Huang, J. Shi, and D. Lin, “Pose guided human video generation,” in ECCV , 2018
2018
Earlier work this paper cites.
H. Cai, C. Bai, Y.-W. Tai, and C.-K. Tang, “Deep video generation, prediction and completion of human action sequences,” in ECCV , 2018
2018
Earlier work this paper cites.
T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, A. Tao, J. Kautz, and B. Catanzaro, “High-resolution image synthesis and semantic manipulation with conditional gans,” in CVPR , 2018
2018
Earlier work this paper cites.
A. Rössler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nießner, “Faceforensics: A large-scale video dataset for forgery detection in human faces,” arXiv , 2018
2018
Earlier work this paper cites.
M. S. Islam, M. S. Rahman, and M. A. Amin, “Beat based realistic dance video generation using deep learning,” in RAAICON , 2019
2019
Earlier work this paper cites.
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe, “First order motion model for image animation,” NeurIPS , 2019
2019
Earlier work this paper cites.
A. Pumarola, J. Sanchez-Riera, G. Choi, A. Sanfeliu, and F. Moreno-Noguer, “3dpeople: Modeling the geometry of dressed humans,” in ICCV , 2019
2019
Earlier work this paper cites.
C. Chan, S. Ginosar, T. Zhou, and A. A. Efros, “Everybody dance now,” in ICCV , 2019
2019
Earlier work this paper cites.
P. Zablotskaia, A. Siarohin, B. Zhao, and L. Sigal, “Dwnet: Dense warp-based network for pose-guided human video generation,” arXiv , 2019
2019
Earlier work this paper cites.
S. Ginosar, A. Bar, G. Kohavi, C. Chan, A. Owens, and J. Malik, “Learning individual styles of conversational gesture,” in CVPR , 2019
2019
Earlier work this paper cites.
Y. Yoon, W.-R. Ko, M. Jang, J. Lee, J. Kim, and G. Lee, “Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots,” in ICRA , 2019
2019
Earlier work this paper cites.
G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black, “Expressive body capture: 3d hands, face, and body from a single image,” in CVPR , 2019
2019
Earlier work this paper cites.
C. Godard, O. Mac Aodha, M. Firman, and G. J. Brostow, “Digging into self-supervised monocular depth estimation,” in ICCV , 2019
2019
Earlier work this paper cites.
L. Liu and G. Feng, “A pilot study on mandarin chinese cued speech,” American Annals of the Deaf , vol. 164, no. 4, pp. 496–518, 2019
2019
Earlier work this paper cites.
L. Yang, Z. Zhao, S. Wang, S. Wang, S. Ma, and W. Gao, “Disentangled human action video generation via decoupled learning,” in ICMEW , 2019
2019
Earlier work this paper cites.
C. Chan, S. Ginosar, T. Zhou, and A. A. Efros, “Everybody dance now,” in ICCV , 2019
2019
Cited alongside, same era.
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe, “First order motion model for image animation,” NeurIPS , 2019
2019
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” NeurIPS , 2020
2020
Cited alongside, same era.
L. Chen, G. Cui, Z. Kou, H. Zheng, and C. Xu, “What comprises a good talking-head video generation?: A survey and benchmark,” arXiv , 2020
2020
Cited alongside, same era.
Y. Yoon, B. Cha, J.-H. Lee, M. Jang, J. Lee, J. Kim, and G. Lee, “Speech gesture generation from the trimodal context of text, audio, and speaker identity,” ACM Transactions on Graphics (TOG) , vol. 39, no. 6, pp. 1–16, 2020
2020
Cited alongside, same era.
Z. Yang, A. Zeng, C. Yuan, and Y. Li, “Effective whole-body pose estimation with two-stages distillation,” in ICCV , 2023
2023
Later among the works it cites.
W. Zhu, X. Ma, Z. Liu, L. Liu, W. Wu, and Y. Wang, “Motionbert: A unified perspective on learning human motion representations,” in ICCV , 2023
2023
Later among the works it cites.
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, “Smpl: A skinned multi-person linear model,” in Seminal Graphics Papers: Pushing the Boundaries, Volume 2 , 2023, pp. 851–866
2023
Later among the works it cites.
R. Jia and S. Pang, “Music2play: Audio-driven instrumental animation,” in CAC , 2023
2023
Later among the works it cites.
Y. Gao, Y. Zhou, J. Wang, X. Li, X. Ming, and Y. Lu, “High-fidelity and freely controllable talking head video generation,” in CVPR , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Luo, J. Ye, R. B. Adams, J. Li, M. G. Newman, and J. Z. Wang, “Arbee: Towards automated recognition of bodily expression of emotion in the wild,” International journal of computer vision , vol. 128, pp. 1–25, 2020
2020
Cited alongside, same era.
J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y. Zhao, D. Liu, Y. Mu, M. Tan, X. Wang et al. , “Deep high-resolution representation learning for visual recognition,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 10, pp. 3349–3364, 2020
2020
Cited alongside, same era.
V. Choutas, G. Pavlakos, T. Bolkart, D. Tzionas, and M. J. Black, “Monocular expressive body regression through body-driven attention,” in ECCV , 2020
2020
Cited alongside, same era.
Z. Teed and J. Deng, “Raft: Recurrent all-pairs field transforms for optical flow,” in ECCV , 2020
2020
Cited alongside, same era.
S. Stoll, S. Hadfield, and R. Bowden, “Signsynth: Data-driven sign language video generation,” in ECCV , 2020
2020
Cited alongside, same era.
M. Liao, S. Zhang, P. Wang, H. Zhu, X. Zuo, and R. Yang, “Speech2video synthesis with 3d skeleton regularization and expressive body poses,” in ACCV , 2020
2020
Cited alongside, same era.
N. Fushishita, A. Tejero-de Pablos, Y. Mukuta, and T. Harada, “Long-term human video generation of multiple futures using poses,” in ECCV , 2020
2020
Cited alongside, same era.
Y. Gan, Z. Yang, X. Yue, L. Sun, and Y. Yang, “Efficient emotional adaptation for audio-driven talking-head generation,” in ICCV , 2023
2023
Later among the works it cites.
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts et al. , “Stable video diffusion: Scaling latent video diffusion models to large datasets,” arXiv , 2023
2023
Later among the works it cites.
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis, “Align your latents: High-resolution video synthesis with latent diffusion models,” in CVPR , 2023
2023
Later among the works it cites.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in ICCV , 2023
2023
Later among the works it cites.
Z. Yang, A. Zeng, C. Yuan, and Y. Li, “Effective whole-body pose estimation with two-stages distillation,” in ICCV , 2023
2023
Later among the works it cites.
J. Karras, A. Holynski, T.-C. Wang, and I. Kemelmacher-Shlizerman, “Dreampose: Fashion image-to-video synthesis via stable diffusion,” in ICCV , 2023
2023
Later among the works it cites.
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, “Smpl: A skinned multi-person linear model,” in Seminal Graphics Papers: Pushing the Boundaries, Volume 2 , 2023
2023
Later among the works it cites.
Y. Wang, X. Ma, X. Chen, C. Chen, A. Dantcheva, B. Dai, and Y. Qiao, “Leo: Generative latent image animator for human video synthesis,” arXiv , 2023
2023
Later among the works it cites.
M. Feng, J. Liu, K. Yu, Y. Yao, Z. Hui, X. Guo, X. Lin, H. Xue, C. Shi, X. Li et al. , “Dreamoving: A human video generation framework based on diffusion models,” arXiv , 2023
2023
Later among the works it cites.
S. F. Bhat, R. Birkl, D. Wofk, P. Wonka, and M. Müller, “Zoedepth: Zero-shot transfer by combining relative and metric depth,” arXiv , 2023
2023
Later among the works it cites.
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu et al. , “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” arXiv , 2023
2023
Later among the works it cites.
W. Lu, Y. Xu, J. Zhang, C. Wang, and D. Tao, “Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpainting,” arXiv , 2023
2023
Later among the works it cites.
X. He, Q. Liu, S. Qian, X. Wang, T. Hu, K. Cao, K. Yan, M. Zhou, and J. Zhang, “Id-animator: Zero-shot identity-preserving human video generation,” arXiv , 2024
2024
Closest in time.
Y. Ma, Y. He, X. Cun, X. Wang, S. Chen, X. Li, and Q. Chen, “Follow your pose: Pose-guided text-to-video generation using pose-free videos,” in AAAI , 2024
2024
Closest in time.
X. He, Q. Huang, Z. Zhang, Z. Lin, Z. Wu, S. Yang, M. Li, Z. Chen, S. Xu, and X. Wu, “Co-speech gesture video generation via motion-decoupled diffusion model,” in CVPR , 2024
2024
Closest in time.
L. Hu, “Animate anyone: Consistent and controllable image-to-video synthesis for character animation,” in CVPR , 2024
2024
Closest in time.
J. Cho, F. D. Puspitasari, S. Zheng, J. Zheng, L.-H. Lee, T.-H. Kim, C. S. Hong, and C. Zhang, “Sora as an agi world model? a complete survey on text-to-video generation,” arXiv , 2024
2024
Closest in time.
C. Li, D. Huang, Z. Lu, Y. Xiao, Q. Pei, and L. Bai, “A survey on long video generation: Challenges, methods, and prospects,” arXiv , 2024
2024
Closest in time.
F. Liao, X. Zou, and W. Wong, “Appearance and pose-guided human generation: A survey,” ACM Computing Surveys , vol. 56, no. 5, pp. 1–35, 2024
2024
Closest in time.
Y. Liu, X. Cun, X. Liu, X. Wang, Y. Zhang, H. Chen, Y. Liu, T. Zeng, R. Chan, and Y. Shan, “Evalcrafter: Benchmarking and evaluating large video generation models,” in CVPR , 2024
2024
Closest in time.
T. Wang, L. Li, K. Lin, Y. Zhai, C.-C. Lin, Z. Yang, H. Zhang, Z. Liu, and L. Wang, “Disco: Disentangled control for realistic human dance generation,” in CVPR , 2024
2024
Closest in time.
L. Liu, L. Liu, and H. Li, “Computation and parameter efficient multi-modal fusion transformer for cued speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
2024
Closest in time.
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, “Depth anything: Unleashing the power of large-scale unlabeled data,” in CVPR , 2024
2024
Closest in time.
T. Kim, C. Kang, J. Park, D. Jeong, C. Yang, S.-J. Kang, and K. Kong, “Human motion aware text-to-video generation with explicit camera control,” in WACV , 2024
2024
Closest in time.
S. Fang, L. Wang, C. Zheng, Y. Tian, and C. Chen, “Signllm: Sign languages production large language models,” arXiv , 2024
2024
Closest in time.
W. Lei, L. Liu, and J. Wang, “Bridge to non-barrier communication: Gloss-prompted fine-grained cued speech gesture generation with diffusion model,” arXiv , 2024
2024
Closest in time.
X. Wang, H. Wang, D. Liu, and W. Cai, “Dance any beat: Blending beats with visuals in dance video generation,” arXiv , 2024
2024
Closest in time.
C. Zhang, C. Wang, Y. Zhao, S. Cheng, L. Luo, and X. Guo, “Dr2: Disentangled recurrent representation learning for data-efficient speech video synthesis,” in WACV , 2024
2024
Closest in time.
S. Hogue, C. Zhang, H. Daruger, Y. Tian, and X. Guo, “Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures,” in CVPR , 2024
2024
Closest in time.
Q. Wang, Z. Jiang, C. Xu, J. Zhang, Y. Wang, X. Zhang, Y. Cao, W. Cao, C. Wang, and Y. Fu, “Vividpose: Advancing stable video diffusion for realistic human image animation,” arXiv , 2024
2024
Closest in time.
S. Tu, Q. Dai, Z. Zhang, S. Xie, Z.-Q. Cheng, C. Luo, X. Han, Z. Wu, and Y.-G. Jiang, “Motionfollower: Editing video motion via lightweight score-guided diffusion,” arXiv , 2024
2024
Closest in time.
Y. Zhang, J. Gu, L.-W. Wang, H. Wang, J. Cheng, Y. Zhu, and F. Zou, “Mimicmotion: High-quality human motion video generation with confidence-aware pose guidance,” arXiv , 2024
2024
Closest in time.
X. Wang, S. Zhang, C. Gao, J. Wang, X. Zhou, Y. Zhang, L. Yan, and N. Sang, “Unianimate: Taming unified video diffusion models for consistent human image animation,” arXiv , 2024
2024
Closest in time.
Z. Xu, J. Zhang, J. H. Liew, H. Yan, J.-W. Liu, C. Zhang, J. Feng, and M. Z. Shou, “Magicanimate: Temporally consistent human image animation using diffusion model,” in CVPR , 2024
2024
Closest in time.
R. Shao, Y. Pang, Z. Zheng, J. Sun, and Y. Liu, “Human4dit: Free-view human video generation with 4d diffusion transformer,” arXiv , 2024
2024
Closest in time.
X. Ma, Y. Wang, G. Jia, X. Chen, Z. Liu, Y.-F. Li, C. Chen, and Y. Qiao, “Latte: Latent diffusion transformer for video generation,” arXiv , 2024
2024
Closest in time.
J. Liu, K. Yu, M. Feng, X. Guo, and M. Cui, “Disentangling foreground and background motion for enhanced realism in human video generation,” arXiv , 2024
2024
Closest in time.
S. Zhu, J. L. Chen, Z. Dai, Y. Xu, X. Cao, Y. Yao, H. Zhu, and S. Zhu, “Champ: Controllable and consistent human image animation with 3d parametric guidance,” arXiv , 2024
2024
Closest in time.
B. Zhu, F. Wang, T. Lu, P. Liu, J. Su, J. Liu, Y. Zhang, Z. Wu, Y.-G. Jiang, and G.-J. Qi, “Poseanimate: Zero-shot high fidelity pose controllable character animation,” arXiv , 2024
2024
Closest in time.
Z. Cai, W. Yin, A. Zeng, C. Wei, Q. Sun, W. Yanjun, H. E. Pang, H. Mei, M. Zhang, L. Zhang et al. , “Smpler-x: Scaling up expressive human pose and shape estimation,” NeurIPS , 2024
2024
Closest in time.
J. Xue, H. Wang, Q. Tian, Y. Ma, A. Wang, Z. Zhao, S. Min, W. Zhao, K. Zhang, H.-Y. Shum et al. , “Follow-your-pose v2: Multiple-condition guided character image animation for stable pose control,” arXiv , 2024
2024
Closest in time.