Fetching the paper…
Reading the bibliography…
Emotional talking head generation has attracted growing attention.
V. Blanz and T. Vetter, “A morphable model for the synthesis of 3d faces,” in Proceedings of the 26th annual conference on Computer graphics and interactive techniques , 1999, pp. 187–194
1999
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.” Journal of machine learning research , vol. 9, no. 11, 2008
2008
Earlier work this paper cites.
N. D. Narvekar and L. J. Karam, “A no-reference perceptual image sharpness metric based on a cumulative probability of blur detection,” in 2009 International Workshop on Quality of Multimedia Experience . IEEE, 2009, pp. 87–91
2009
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in ICML , 2015
2015
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in Asian conference on computer vision . Springer, 2016, pp. 251–263
2016
Earlier work this paper cites.
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Synthesizing obama: learning lip sync from audio,” ACM Transactions on Graphics (ToG) , vol. 36, no. 4, pp. 1–13, 2017
2017
Earlier work this paper cites.
J. S. Chung, A. Jamaludin, and A. Zisserman, “You said that?” arXiv preprint arXiv:1705.02966 , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30. Curran Associates, Inc., 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
L. Chen, Z. Li, R. K. Maddox, Z. Duan, and C. Xu, “Lip movements generation at a glance,” in ECCV , 2018, pp. 520–535
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” PloS one , vol. 13, no. 5, p. e0196391, 2018
2018
Earlier work this paper cites.
O. Fried, A. Tewari, M. Zollhöfer, A. Finkelstein, E. Shechtman, D. B. Goldman, K. Genova, Z. Jin, C. Theobalt, and M. Agrawala, “Text-based editing of talking-head video,” ACM Transactions on Graphics (TOG) , vol. 38, no. 4, pp. 1–14, 2019
2019
Earlier work this paper cites.
N. Sadoughi and C. Busso, “Speech-driven expressive talking lips with conditional sequential generative adversarial networks,” IEEE Transactions on Affective Computing , vol. 12, no. 4, pp. 1031–1044, 2019
2019
Earlier work this paper cites.
K. Vougioukas, S. Petridis, and M. Pantic, “Realistic speech-driven facial animation with gans,” International Journal of Computer Vision , pp. 1–16, 2019
2019
Earlier work this paper cites.
H. Zhou, Y. Liu, Z. Liu, P. Luo, and X. Wang, “Talking face generation by adversarially disentangled audio-visual representation,” in AAAI , vol. 33, no. 01, 2019, pp. 9299–9306
2019
Earlier work this paper cites.
L. Chen, R. K. Maddox, Z. Duan, and C. Xu, “Hierarchical cross-modal talking face generation with dynamic pixel-wise loss,” in CVPR , 2019, pp. 7832–7841
2019
Earlier work this paper cites.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in CVPR , 2019, pp. 4690–4699
2019
Earlier work this paper cites.
D. Cudeiro, T. Bolkart, C. Laidlaw, A. Ranjan, and M. J. Black, “Capture, learning, and synthesis of 3d speaking styles,” in CVPR , 2019, pp. 10 101–10 111
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 6840–6851
2020
Earlier work this paper cites.
D. Das, S. Biswas, S. Sinha, and B. Bhowmick, “Speech-driven facial animation using cascaded gans for learning of motion and texture,” in ECCV . Springer, 2020, pp. 408–424
2020
Earlier work this paper cites.
K. Wang, Q. Wu, L. Song, Z. Yang, W. Wu, C. Qian, R. He, Y. Qiao, and C. C. Loy, “Mead: A large-scale audio-visual dataset for emotional talking-face generation,” in ECCV . Springer, 2020, pp. 700–717
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Thies, M. Elgharib, A. Tewari, C. Theobalt, and M. Nießner, “Neural voice puppetry: Audio-driven facial reenactment,” in ECCV . Springer, 2020, pp. 716–731
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 484–492
2020
Earlier work this paper cites.
Y. Zhou, X. Han, E. Shechtman, J. Echevarria, E. Kalogerakis, and D. Li, “Makelttalk: speaker-aware talking-head animation,” ACM Transactions on Graphics (TOG) , vol. 39, no. 6, pp. 1–15, 2020
2020
Earlier work this paper cites.
L. Chen, G. Cui, C. Liu, Z. Li, Z. Kou, Y. Xu, and C. Xu, “Talking-head generation with rhythmic head motion,” in ECCV . Springer, 2020, pp. 35–51
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” NeurIPS , 2021
2021
Earlier work this paper cites.
X. Ji, H. Zhou, K. Wang, W. Wu, C. C. Loy, X. Cao, and F. Xu, “Audio-driven emotional video portraits,” in CVPR , 2021, pp. 14 080–14 089
2021
Earlier work this paper cites.
Y. Lu, J. Chai, and X. Cao, “Live speech portraits: real-time photorealistic talking-head animation,” ACM Transactions on Graphics (TOG) , vol. 40, no. 6, pp. 1–17, 2021
2021
Earlier work this paper cites.
A. Lahiri, V. Kwatra, C. Frueh, J. Lewis, and C. Bregler, “Lipsync3d: Data-efficient learning of personalized 3d talking faces from video using pose and lighting normalization,” in CVPR , 2021, pp. 2755–2764
2021
Earlier work this paper cites.
C. Zhang, S. Ni, Z. Fan, H. Li, M. Zeng, M. Budagavi, and X. Guo, “3d talking face with personalized pose dynamics,” IEEE Transactions on Visualization and Computer Graphics , 2021
2021
Cited alongside, same era.
C. Zhang, Y. Zhao, Y. Huang, M. Zeng, S. Ni, M. Budagavi, and X. Guo, “Facial: Synthesizing dynamic talking face with implicit attribute learning,” in ICCV , 2021, pp. 3867–3876
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Z. Zhang, L. Li, Y. Ding, and C. Fan, “Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,” in CVPR , 2021, pp. 3661–3670
2021
Cited alongside, same era.
J. Wang, K. Zhao, Y. Ma, S. Zhang, Y. Zhang, Y. Shen, D. Zhao, and J. Zhou, “Facecomposer: A unified model for versatile facial content creation,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Closest in time.
2023
Closest in time.
W. Zhang, X. Cun, X. Wang, Y. Zhang, X. Shen, Y. Guo, Y. Shan, and F. Wang, “Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 8652–8661
2023
Closest in time.
Z. Sheng, L. Nie, M. Liu, Y. Wei, and Z. Gao, “Towards fine-grained talking face generation,” IEEE Transactions on Image Processing , 2023
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
H. Zhou, Y. Sun, W. Wu, C. C. Loy, X. Wang, and Z. Liu, “Pose-controllable talking face generation by implicitly modularized audio-visual representation,” in CVPR , 2021, pp. 4176–4186
2021
Cited alongside, same era.
M. Cao, H. Huang, H. Wang, X. Wang, L. Shen, S. Wang, L. Bao, Z. Li, and J. Luo, “Unifacegan: a unified framework for temporally consistent facial video editing,” IEEE Transactions on Image Processing , vol. 30, pp. 6107–6116, 2021
2021
Cited alongside, same era.
X. Wu, Q. Zhang, Y. Wu, H. Wang, S. Li, L. Sun, and X. Li, “F 3 a-gan: Facial flow for face animation with generative adversarial networks,” IEEE Transactions on Image Processing , vol. 30, pp. 8658–8670, 2021
2021
Cited alongside, same era.
Y. Ren, G. Li, Y. Chen, T. H. Li, and S. Liu, “Pirenderer: Controllable portrait image generation via semantic neural rendering,” in ICCV , 2021, pp. 13 759–13 768
2021
Cited alongside, same era.
B. Liang, Y. Pan, Z. Guo, H. Zhou, Z. Hong, X. Han, J. Han, J. Liu, E. Ding, and J. Wang, “Expressive talking head generation with granular audio-visual control,” in CVPR , 2022, pp. 3387–3396
2022
Cited alongside, same era.
2022
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR , 2022
2022
Cited alongside, same era.
2023
Closest in time.
S. Gururani, A. Mallya, T.-C. Wang, R. Valle, and M.-Y. Liu, “Space: Speech-driven portrait animation with controllable expression,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 20 914–20 923
2023
Closest in time.
S. Tan, B. Ji, and Y. Pan, “Emmn: Emotional motion memory network for audio-driven emotional talking face generation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 22 146–22 156
2023
Closest in time.
Z. Peng, H. Wu, Z. Song, H. Xu, X. Zhu, J. He, H. Liu, and Z. Fan, “Emotalk: Speech-driven emotional disentanglement for 3d face animation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 20 687–20 697
2023
Closest in time.
Y. Zhang, X. Xu, Y. Zhao, Y. Wen, Z. Tang, and M. Liu, “Facial prior guided micro-expression generation,” IEEE Transactions on Image Processing , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
S. Wang, Y. Ma, and Y. Ding, “Exploring complementary features in multi-modal speech emotion recognition,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Closest in time.
S. Wang, Y. Ma, Y. Ding, Z. Hu, C. Fan, T. Lv, Z. Deng, and X. Yu, “Styletalk++: A unified framework for controlling the speaking styles of talking heads,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Peng, W. Hu, Y. Shi, X. Zhu, X. Zhang, H. Zhao, J. He, H. Liu, and Z. Fan, “Synctalk: The devil is in the synchronization for talking head synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 666–676
2024
Closest in time.
2024
Closest in time.
X. Gao, Y. Yang, Y. Wu, S. Du, and G.-J. Qi, “Multi-condition latent diffusion network for scene-aware neural human motion prediction,” IEEE Transactions on Image Processing , 2024
2024
Closest in time.
Y. Wang, H. Liu, Y. Feng, Z. Li, X. Wu, and C. Zhu, “Headdiff: Exploring rotation uncertainty with diffusion models for head pose estimation,” IEEE Transactions on Image Processing , 2024
2024
Closest in time.
S. Welker, H. N. Chapman, and T. Gerkmann, “Driftrec: Adapting diffusion models to blind jpeg restoration,” IEEE Transactions on Image Processing , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.