Fetching the paper…
Reading the bibliography…
Different people speak with diverse personalized speaking styles.
Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation
Li, L.; Wang, S.; Zhang, Z.; Ding, Y.; Zheng, Y.; Yu, X.; and Fan, C. 2021 · 1920
Earlier work this paper cites.
A morphable model for the synthesis of 3D faces
Blanz, V.; and Vetter, T. 1999 · 1999
Earlier work this paper cites.
Everybody’s talkin’: Let me talk as you want
Song, L.; Wu, W.; Qian, C.; He, R.; and Loy, C. C. 2020 · 2001
Earlier work this paper cites.
Audio-driven talking face video generation with learning-based personalized head pose
Yi, R.; Ye, Z.; Zhang, J.; Bao, H.; and Liu, Y.-J. 2020 · 2002
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004 · 2004
Earlier work this paper cites.
What comprises a good talking-head video generation?: A survey and benchmark
Chen, L.; Cui, G.; Kou, Z.; Zheng, H.; and Xu, C. 2020a · 2005
Earlier work this paper cites.
Self-attention encoding and pooling for speaker recognition
Safari, P.; India, M.; and Hernando, J. 2020 · 2008
Earlier work this paper cites.
Visualizing data using t-SNE
Van der Maaten, L.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
A no-reference perceptual image sharpness metric based on a cumulative probability of blur detection
Narvekar, N. D.; and Karam, L. J. 2009 · 2009
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Chung, J. S.; and Zisserman, A. 2016 · 2016
Earlier work this paper cites.
Chung, J. S.; Jamaludin, A.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Isola, P.; Zhu, J.-Y.; Zhou, T.; and Efros, A. A. 2017 · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Qi, C. R.; Su, H.; Mo, K.; and Guibas, L. J. 2017 · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Suwajanakorn, S.; Seitz, S. M.; and Kemelmacher-Shlizerman, I. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Lip movements generation at a glance
Chen, L.; Li, Z.; Maddox, R. K.; Duan, Z.; and Xu, C. 2018 · 2018
Earlier work this paper cites.
Triplet loss in siamese network for object tracking
Dong, X.; and Shen, J. 2018 · 2018
Cited alongside, same era.
X-vectors: Robust dnn embeddings for speaker recognition
Snyder, D.; Garcia-Romero, D.; Sell, G.; Povey, D.; and Khudanpur, S. 2018 · 2018
Cited alongside, same era.
Talking face generation by conditional recurrent adversarial network
Song, Y.; Zhu, J.; Li, D.; Wang, X.; and Qi, H. 2018 · 2018
Cited alongside, same era.
X2face: A network for controlling face generation using images, audio, and pose codes
Wiles, O.; Koepke, A.; and Zisserman, A. 2018 · 2018
Cited alongside, same era.
Face super-resolution guided by facial component heatmaps
Yu, X.; Fernando, B.; Ghanem, B.; Porikli, F.; and Hartley, R. 2018 · 2018
Cited alongside, same era.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Rethinking the value of transformer components
Wang, W.; and Tu, Z. 2020 · 2020
Later among the works it cites.
MakeltTalk: speaker-aware talking-head animation
Zhou, Y.; Han, X.; Shechtman, E.; Echevarria, J.; Kalogerakis, E.; and Li, D. 2020 · 2020
Later among the works it cites.
AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis
Guo, Y.; Chen, K.; Liang, S.; Liu, Y.; Bao, H.; and Zhang, J. 2021 · 2021
Later among the works it cites.
Audio-driven emotional video portraits
Ji, X.; Zhou, H.; Wang, K.; Wu, W.; Loy, C. C.; Cao, X.; and Xu, F. 2021 · 2021
Later among the works it cites.
LipSync3D: Data-Efficient Learning of Personalized 3D Talking Faces from Video using Pose and Lighting Normalization
Lahiri, A.; Kwatra, V.; Frueh, C.; Lewis, J.; and Bregler, C. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, L.; Maddox, R. K.; Duan, Z.; and Xu, C. 2019 · 2019
Cited alongside, same era.
Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set
Deng, Y.; Yang, J.; Xu, S.; Chen, D.; Jia, Y.; and Tong, X. 2019 · 2019
Cited alongside, same era.
Text-based editing of talking-head video
Fried, O.; Tewari, A.; Zollhöfer, M.; Finkelstein, A.; Shechtman, E.; Goldman, D. B.; Genova, K.; Jin, Z.; Theobalt, C.; and Agrawala, M. 2019 · 2019
Cited alongside, same era.
Speech-driven expressive talking lips with conditional sequential generative adversarial networks
Sadoughi, N.; and Busso, C. 2019 · 2019
Cited alongside, same era.
First order motion model for image animation
Siarohin, A.; Lathuilière, S.; Tulyakov, S.; Ricci, E.; and Sebe, N. 2019 · 2019
Cited alongside, same era.
Realistic speech-driven facial animation with gans
Vougioukas, K.; Petridis, S.; and Pantic, M. 2019 · 2019
Cited alongside, same era.
Condconv: Conditionally parameterized convolutions for efficient inference
Yang, B.; Bender, G.; Le, Q. V.; and Ngiam, J. 2019 · 2019
Cited alongside, same era.
Qian, S.; Tu, Z.; Zhi, Y.; Liu, W.; and Gao, S. 2021 · 2021
Later among the works it cites.
Pirenderer: Controllable portrait image generation via semantic neural rendering
Ren, Y.; Li, G.; Chen, Y.; Li, T. H.; and Liu, S. 2021 · 2021
Later among the works it cites.
Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion
Wang, S.; Li, L.; Ding, Y.; Fan, C.; and Yu, X. 2021 · 2021
Later among the works it cites.
Imitating arbitrary talking style for realistic audio-driven talking face synthesis
Wu, H.; Jia, J.; Wang, H.; Dou, Y.; Duan, C.; and Deng, Q. 2021 · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Zhou, H.; Sun, Y.; Wu, W.; Loy, C. C.; Wang, X.; and Liu, Z. 2021 · 2021
Later among the works it cites.
Deep audio-visual learning: A survey
Zhu, H.; Luo, M.-D.; Wang, R.; Zheng, A.-H.; and He, R. 2021 · 2021
Later among the works it cites.
EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion Model
Ji, X.; Zhou, H.; Wang, K.; Wu, Q.; Wu, W.; Xu, F.; and Cao, X. 2022 · 2022
Later among the works it cites.
Expressive talking head generation with granular audio-visual control
Liang, B.; Pan, Y.; Guo, Z.; Zhou, H.; Hong, Z.; Han, X.; Han, J.; Liu, J.; Ding, E.; and Wang, J. 2022 · 2022
Later among the works it cites.
Semantic-aware implicit neural audio-driven video portrait generation
Liu, X.; Xu, Y.; Wu, Q.; Zhou, H.; Wu, W.; and Zhou, B. 2022 · 2022
Later among the works it cites.
Emotion-Controllable Generalized Talking Face Generation
Sinha, S.; Biswas, S.; Yadav, R.; and Bhowmick, B. 2022 · 2022
Later among the works it cites.
One-shot talking face generation from single-speaker audio-visual correlation learning
Wang, S.; Li, L.; Ding, Y.; and Yu, X. 2022 · 2022
Later among the works it cites.
Styleheat: One-shot high-resolution editable talking face generation via pretrained stylegan
Yin, F.; Zhang, Y.; Cun, X.; Cao, M.; Fan, Y.; Wang, X.; Bai, Q.; Wu, B.; Wang, J.; and Yang, Y. 2022 · 2022
Later among the works it cites.