Fetching the paper…
Reading the bibliography…
Audio-driven talking head generation has drawn growing attention.
L. Li, S. Wang, Z. Zhang, Y. Ding, Y. Zheng, X. Yu, and C. Fan, “Write-a-speaker: Text-based emotional and rhythmic talking-head generation,” in AAAI , vol. 35, no. 3, 2021, pp. 1911–1920
1920
Earlier work this paper cites.
E. Cosatto and H. Graf, “Photo-realistic talking-heads from image samples,” IEEE Transactions on Multimedia , vol. 2, no. 3, pp. 152–163, 2000
2000
Earlier work this paper cites.
P. Ekman, “Facial action coding system (facs),” A Human Face, Salt Lake City , 2002
2002
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE TIP , 2004
2004
Earlier work this paper cites.
K.-H. Choi and J.-N. Hwang, “Automatic creation of a talking head from a video sequence,” IEEE Transactions on Multimedia , vol. 7, no. 4, pp. 628–637, 2005
2005
Earlier work this paper cites.
N. D. Narvekar and L. J. Karam, “A no-reference perceptual image sharpness metric based on a cumulative probability of blur detection,” in 2009 International Workshop on Quality of Multimedia Experience . IEEE, 2009, pp. 87–91
2009
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in ACCV . Springer, 2016, pp. 251–263
2016
Earlier work this paper cites.
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Synthesizing obama: learning lip sync from audio,” ToG , vol. 36, no. 4, pp. 1–13, 2017
2017
Earlier work this paper cites.
J. S. Chung, A. Jamaludin, and A. Zisserman, “You said that?” arXiv preprint arXiv:1705.02966 , 2017
2017
Earlier work this paper cites.
L. Chen, Z. Li, R. K. Maddox, Z. Duan, and C. Xu, “Lip movements generation at a glance,” in ECCV , 2018, pp. 520–535
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
H. Zhou, Y. Liu, Z. Liu, P. Luo, and X. Wang, “Talking face generation by adversarially disentangled audio-visual representation,” in AAAI , vol. 33, no. 01, 2019, pp. 9299–9306
2019
Earlier work this paper cites.
L. Chen, R. K. Maddox, Z. Duan, and C. Xu, “Hierarchical cross-modal talking face generation with dynamic pixel-wise loss,” in CVPR , 2019, pp. 7832–7841
2019
Earlier work this paper cites.
N. Sadoughi and C. Busso, “Speech-driven expressive talking lips with conditional sequential generative adversarial networks,” TAC , vol. 12, no. 4, pp. 1031–1044, 2019
2019
Earlier work this paper cites.
K. Vougioukas, S. Petridis, and M. Pantic, “Realistic speech-driven facial animation with gans,” IJCV , 2019
2019
Earlier work this paper cites.
K. Wang, Q. Wu, L. Song, Z. Yang, W. Wu, C. Qian, R. He, Y. Qiao, and C. C. Loy, “Mead: A large-scale audio-visual dataset for emotional talking-face generation,” in ECCV . Springer, 2020, pp. 700–717
2020
Earlier work this paper cites.
J. Thies, M. Elgharib, A. Tewari, C. Theobalt, and M. Nießner, “Neural voice puppetry: Audio-driven facial reenactment,” in ECCV . Springer, 2020, pp. 716–731
2020
Earlier work this paper cites.
K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in ACMMM , 2020, pp. 484–492
2020
Earlier work this paper cites.
L. Chen, G. Cui, C. Liu, Z. Li, Z. Kou, Y. Xu, and C. Xu, “Talking-head generation with rhythmic head motion,” in ECCV , 2020
2020
Earlier work this paper cites.
Y. Zhou, X. Han, E. Shechtman, J. Echevarria, E. Kalogerakis, and D. Li, “Makelttalk: speaker-aware talking-head animation,” TOG , 2020
2020
Earlier work this paper cites.
J. Hong, H. J. Lee, Y. Kim, and Y. M. Ro, “Face tells detailed expression: Generating comprehensive facial expression sentence through facial action units,” in International Conference on Multimedia Modeling . Springer, 2020, pp. 100–111
2020
Earlier work this paper cites.
X. Ji, H. Zhou, K. Wang, W. Wu, C. C. Loy, X. Cao, and F. Xu, “Audio-driven emotional video portraits,” in CVPR , 2021, pp. 14 080–14 089
2021
Cited alongside, same era.
C. Zhang, S. Ni, Z. Fan, H. Li, M. Zeng, M. Budagavi, and X. Guo, “3d talking face with personalized pose dynamics,” TVCG , 2021
2021
Cited alongside, same era.
C. Zhang, Y. Zhao, Y. Huang, M. Zeng, S. Ni, M. Budagavi, and X. Guo, “Facial: Synthesizing dynamic talking face with implicit attribute learning,” in ICCV , 2021, pp. 3867–3876
2021
Cited alongside, same era.
Y. Guo, K. Chen, S. Liang, Y. Liu, H. Bao, and J. Zhang, “Ad-nerf: Audio driven neural radiance fields for talking head synthesis,” ICCV , 2021
2021
Cited alongside, same era.
Z. Zhang, L. Li, Y. Ding, and C. Fan, “Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,” in CVPR , 2021, pp. 3661–3670
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR , 2022, pp. 10 684–10 695
2022
Later among the works it cites.
J. Sun, Q. Deng, Q. Li, M. Sun, M. Ren, and Z. Sun, “Anyface: Free-style text-to-face synthesis and manipulation,” in CVPR , 2022, pp. 18 687–18 696
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
S. Wang, L. Li, Y. Ding, C. Fan, and X. Yu, “Audio2head: Audio-driven one-shot talking-head generation with natural head motion,” IJCAI , 2021
2021
Cited alongside, same era.
H. Zhou, Y. Sun, W. Wu, C. C. Loy, X. Wang, and Z. Liu, “Pose-controllable talking face generation by implicitly modularized audio-visual representation,” in CVPR , 2021, pp. 4176–4186
2021
Cited alongside, same era.
H. Wu, J. Jia, H. Wang, Y. Dou, C. Duan, and Q. Deng, “Imitating arbitrary talking style for realistic audio-driven talking face synthesis,” in ACMMM , 2021, pp. 1478–1486
2021
Cited alongside, same era.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in ICML . PMLR, 2021, pp. 8821–8831
2021
Cited alongside, same era.
2021
Cited alongside, same era.
A. Ghosh, N. Cheema, C. Oguz, C. Theobalt, and P. Slusallek, “Synthesis of compositional animations from textual descriptions,” in ICCV , 2021, pp. 1396–1406
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML . PMLR, 2021, pp. 8748–8763
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
K. Chen, X. Yang, C. Fan, W. Zhang, and Y. Ding, “Semantic-rich facial emotional expression recognition,” IEEE TAC , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Ma, S. Wang, Z. Hu, C. Fan, T. Lv, Y. Ding, Z. Deng, and X. Yu, “Styletalk: One-shot talking head generation with controllable speaking styles,” in AAAI , 2023
2023
Closest in time.
J. Yu, H. Zhu, L. Jiang, C. C. Loy, W. Cai, and W. Wu, “CelebV-Text: A large-scale facial text-video dataset,” in CVPR , 2023
2023
Closest in time.
X. Wang, Q. Xie, J. Zhu, L. Xie, and O. Scharenborg, “Anyonenet: Synchronized speech and talking head generation for arbitrary persons,” IEEE Transactions on Multimedia , vol. 25, pp. 6717–6728, 2023
2023
Closest in time.
Z. Yu, Z. Yin, D. Zhou, D. Wang, F. Wong, and B. Wang, “Talking head generation with probabilistic audio-to-visual diffusion priors,” ICCV , 2023
2023
Closest in time.
Z. Ye, M. Xia, R. Yi, J. Zhang, Y.-K. Lai, X. Huang, G. Zhang, and Y.-J. Liu, “Audio-driven talking face video generation with dynamic convolution kernels,” IEEE Transactions on Multimedia , vol. 25, pp. 2033–2046, 2023
2023
Closest in time.
W. Zhang, X. Cun, X. Wang, Y. Zhang, X. Shen, Y. Guo, Y. Shan, and F. Wang, “Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,” in CVPR , 2023, pp. 8652–8661
2023
Closest in time.
C. Xu, J. Zhu, J. Zhang, Y. Han, W. Chu, Y. Tai, C. Wang, Z. Xie, and Y. Liu, “High-fidelity generalized emotional talking face generation with multi-modal emotion space learning,” in CVPR , 2023
2023
Closest in time.
Y. Gan, Z. Yang, X. Yue, L. Sun, and Y. Yang, “Efficient emotional adaptation for audio-driven talking-head generation,” in ICCV , 2023, pp. 22 634–22 645
2023
Closest in time.
D. Wang, Y. Deng, Z. Yin, H.-Y. Shum, and B. Wang, “Progressive disentangled representation learning for fine-grained controllable talking head synthesis,” in CVPR , 2023, pp. 17 979–17 989
2023
Closest in time.
S. Shen, W. Li, X. Huang, Z. Zhu, J. Zhou, and J. Lu, “Sd-nerf: Towards lifelike talking head animation via spatially-adaptive dual-driven nerfs,” IEEE Transactions on Multimedia , vol. 26, pp. 3221–3234, 2024
2024
Closest in time.
X. Qi, C. Liu, L. Li, J. Hou, H. Xin, and X. Yu, “Emotiongesture: Audio-driven diverse emotional co-speech 3d gesture generation,” IEEE Transactions on Multimedia , 2024
2024
Closest in time.