Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures.
Animated conversation: rule-based generation of facial expression, gesture & spoken intonation for multiple conversational agents
Cassell, J., Pelachaud, C., Badler, N., Steedman, M., Achorn, B., Becket, T., Douville, B., Prevost, S., and Stone, M · 1994
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Wang, Z., Bovik, A., Sheikh, H., and Simoncelli, E · 2003
Earlier work this paper cites.
Gesture generation by imitation: from human behavior to computer character animation
Kipp, M · 2005
Earlier work this paper cites.
Biomechanics and motor control of human movement
Winter, D. A · 2009
Earlier work this paper cites.
Image quality metrics: Psnr vs. ssim
Horé, A. and Ziou, D · 2010
Earlier work this paper cites.
Gesture and speech in interaction: An overview
Wagner, P., Malisz, Z., and Kopp, S · 2013
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Chung, J. S. and Zisserman, A · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Embodied hands: Modeling and capturing hands and bodies together
Romero, J., Tzionas, D., and Black, M. J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W. T., and Rubinstein, M · 2018
Earlier work this paper cites.
Relaxedik: Real-time synthesis of accurate and feasible robot arm motion
Rakita, D., Mutlu, B., and Gleicher, M · 2018
Earlier work this paper cites.
Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph
Zadeh, A. B., Liang, P. P., Poria, S., Cambria, E., and Morency, L.-P · 2018
Earlier work this paper cites.
Expressive body capture: 3D hands, face, and body from a single image
Pavlakos, G., Choutas, V., Ghorbani, N., Bolkart, T., Osman, A. A. A., Tzionas, D., and Black, M. J · 2019
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
Schneider, S., Baevski, A., Collobert, R., and Auli, M · 2019
Cited alongside, same era.
Fvd: A new metric for video generation
Unterthiner, T., van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2019
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Learning speech-driven 3d conversational gestures from video, 2021
Habibie, I., Xu, W., Mehta, D., Liu, L., Seidel, H.-P., Pons-Moll, G., Elgharib, M., and Theobalt, C · 2021
Cited alongside, same era.
Nmpc-mp: Real-time nonlinear model predictive control for safe motion planning in manipulator teleoperation
Hu, S., Babaians, E., Karimi, M., and Steinbach, E · 2021
Cited alongside, same era.
Chen, J., Liu, Y., Wang, J., Zeng, A., Li, Y., and Chen, Q · 2024
Later among the works it cites.
Vlogger: Multimodal diffusion for embodied avatar synthesis, 2024
Corona, E., Zanfir, A., Bazavan, E. G., Kolotouros, N., Alldieck, T., and Sminchisescu, C · 2024
Later among the works it cites.
Ltx-video: Realtime video latent diffusion
HaCohen, Y., Chiprut, N., Brazowski, B., Shalem, D., Moshe, D., Richardson, E., Levin, E., Shiran, G., Zabari, N., Gordon, O., Panet, P., Weissbuch, S., Kulikov, V., Bitterman, Y., Melumian, Z., and Bibi, O · 2024
Later among the works it cites.
Co-speech gesture video generation via motion-decoupled diffusion model, 2024
He, X., Huang, Q., Zhang, Z., Lin, Z., Wu, Z., Yang, S., Li, M., Chen, Z., Xu, S., and Wu, X · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Vote for grasp poses from noisy point sets by learning from human
Tian, L., Wu, J., Xiong, Z., and Zhu, X · 2021
Cited alongside, same era.
Hit-dvae: Human motion generation via hierarchical transformer dynamical vae
Bie, X., Guo, W., Leglaive, S., Girin, L., Moreno-Noguer, F., and Alameda-Pineda, X · 2022
Cited alongside, same era.
Faceformer: Speech-driven 3d facial animation with transformers
Fan, Y., Lin, Z., Saito, J., Wang, W., and Komura, T · 2022
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Chen, J., Yu, J., Ge, C., Yao, L., Xie, E., Wu, Y., Wang, Z., Kwok, J., Luo, P., Lu, H., and Li, Z · 2023
Cited alongside, same era.
Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures
Hogue, S., Zhang, C., Daruger, H., Tian, Y., and Guo, X · 2024
Later among the works it cites.
Advmt: Adversarial motion transformer for long-term human motion prediction
Idrees, S., Choi, J., and Sohn, S · 2024
Later among the works it cites.
Loopy: Taming audio-driven portrait avatar with long-term motion dependency, 2024
Jiang, J., Liang, C., Yang, J., Lin, G., Zhong, T., and Zheng, Y · 2024
Later among the works it cites.
Cyberhost: Taming audio-driven avatar diffusion model with region codebook attention, 2024
Lin, G., Jiang, J., Liang, C., Zhong, T., Yang, J., and Zheng, Y · 2024
Later among the works it cites.
Echomimicv2: Towards striking, simplified, and semi-body human animation, 2024
Meng, R., Zhang, X., Li, Y., and Ma, C · 2024
Later among the works it cites.
Cocogesture: Toward coherent co-speech 3d gesture generation in the wild, 2024
Qi, X., Zhang, H., Wang, Y., Pan, J., Liu, C., Li, P., Chi, X., Li, M., Xue, W., Zhang, S., Luo, W., Liu, Q., and Guo, Y · 2024
Later among the works it cites.
Hallo: Hierarchical audio-driven visual synthesis for portrait image animation, 2024
Xu, M., Li, H., Su, Q., Shang, H., Zhang, L., and Liu, C · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., et al · 2024
Later among the works it cites.
Mimicmotion: High-quality human motion video generation with confidence-aware pose guidance
Zhang, Y., Gu, J., Wang, L.-W., Wang, H., Cheng, J., Zhu, Y., and Zou, F · 2024
Later among the works it cites.
Hunyuanvideo: A systematic framework for large video generative models, 2025
Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., Wu, K., Lin, Q., Yuan, J., Long, Y., Wang, A., Wang, A., Li, C., Huang, D., Yang, F., Tan, H., Wang, H., Song, J., Bai, J., Wu, J., Xue, J., Wang, J., Wang, K., Liu, M., Li, P., Li, S., Wang, W., Yu, W., Deng, X., Li, Y., Chen, Y., Cui, Y., Peng, Y., Yu, Z., He, Z., Xu, Z., Zhou, Z., Xu, Z., Tao, Y., Lu, Q., Liu, S., Zhou, D., Wang, H., Yang, Y., Wang, D., Liu, Y., Jiang, J., and Zhong, C · 2025
Closest in time.
Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions
Tian, L., Wang, Q., Zhang, B., and Bo, L · 2025
Closest in time.