Fetching the paper…
Reading the bibliography…
Achieving disentangled control over multiple facial motions and accommodating diverse input modalities greatly enhances the application and entertainment of the talking head generation.
Facial action coding system
Paul Ekman and Wallace V Friesen · 1978
Earlier work this paper cites.
A morphable model for the synthesis of 3d faces
Volker Blanz and Thomas Vetter · 1999
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli · 2004
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Variational inference with normalizing flows
Danilo Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Lip reading in the wild
Joon Son Chung and Andrew Zisserman · 2017
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization
Xun Huang and Serge Belongie · 2017
Earlier work this paper cites.
Deep audio-visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2018
Earlier work this paper cites.
Lip movements generation at a glance
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Earlier work this paper cites.
Deepfake video detection using recurrent neural networks
David Guera and Edward J. Delp · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Earlier work this paper cites.
Frame attention networks for facial expression recognition in videos
Debin Meng, Xiaojiang Peng, Kai Wang, and Yu Qiao · 2019
Earlier work this paper cites.
First order motion model for image animation
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe · 2019
Earlier work this paper cites.
Talking face generation by conditional recurrent adversarial network
Yang Song, Jingwen Zhu, Dawei Li, Andy Wang, and Hairong Qi · 2019
Earlier work this paper cites.
Few-shot adversarial learning of realistic neural talking head models
Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky · 2019
Earlier work this paper cites.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Earlier work this paper cites.
What makes fake images detectable? understanding properties that generalize
Lucy Chai, David Bau, Ser-Nam Lim, and Phillip Isola · 2020
Earlier work this paper cites.
Talking-head generation with rhythmic head motion
Lele Chen, Guofeng Cui, Celong Liu, Zhong Li, Ziyi Kou, Yi Xu, and Chenliang Xu · 2020
Earlier work this paper cites.
Speech-driven facial animation using cascaded gans for learning of motion and texture
Dipanjan Das, Sandika Biswas, Sanjana Sinha, and Brojeshwar Bhowmick · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Earlier work this paper cites.
Supervised contrastive learning
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan · 2020
Earlier work this paper cites.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Earlier work this paper cites.
pytorch-fid: FID Score for PyTorch
Maximilian Seitzer · 2020
Cited alongside, same era.
Neural voice puppetry: Audio-driven facial reenactment
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner · 2020
Cited alongside, same era.
Mead: A large-scale audio-visual dataset for emotional talking-face generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy · 2020
Cited alongside, same era.
Makelttalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li · 2020
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Cited alongside, same era.
Audio-driven emotional video portraits
One-shot talking face generation from single-speaker audio-visual correlation learning
Suzhen Wang, Lincheng Li, Yu Ding, and Xin Yu · 2022
Later among the works it cites.
Face2face
Kewei Yang, Kang Chen, Daoliang Guo, Song-Hai Zhang, Yuan-Chen Guo, and Weidong Zhang · 2022
Later among the works it cites.
Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan
Fei Yin, Yong Zhang, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang, Qingyan Bai, Baoyuan Wu, Jue Wang, and Yujiu Yang · 2022
Later among the works it cites.
Video rewrite: Driving visual speech with audio
Christoph Bregler, Michele Covell, and Malcolm Slaney · 2023
Later among the works it cites.
Vast: Vivify your talking avatar via zero-shot expressive facial style transfer
Liyang Chen, Zhiyong Wu, Runnan Li, Weihong Bao, Jun Ling, Xu Tan, and Sheng Zhao · 2023
Later among the works it cites.
Efficient emotional adaptation for audio-driven talking-head generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu · 2021
Cited alongside, same era.
Emoberta: Speaker-aware emotion recognition in conversation with roberta
Taewoon Kim and Piek Vossen · 2021
Cited alongside, same era.
Ai-generated characters for supporting personalized learning and well-being
Pat Pataranutaporn, Valdemar Danry, Joanne Leong, Parinya Punpongsanon, Dan Novy, Pattie Maes, and Misha Sra · 2021
Cited alongside, same era.
Pirenderer: Controllable portrait image generation via semantic neural rendering
Yurui Ren, Ge Li, Yuanqi Chen, Thomas H Li, and Shan Liu · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Cited alongside, same era.
Emotion-controllable generalized talking face generation
Sanjana Sinha, Sandika Biswas, Ravindra Yadav, and Brojeshwar Bhowmick · 2021
Cited alongside, same era.
Audio2head: Audio-driven one-shot talking-head generation with natural head motion
S Wang, L Li, Y Ding, C Fan, and X Yu · 2021
Cited alongside, same era.
Yuan Gan, Zongxin Yang, Xihang Yue, Lingyun Sun, and Yi Yang · 2023
Later among the works it cites.
Emotionally enhanced talking face generation
Sahil Goyal, Sarthak Bhagat, Shagun Uppal, Hitkul Jangra, Yi Yu, Yifang Yin, and Rajiv Ratn Shah · 2023
Later among the works it cites.
Implicit identity representation conditioned memory compensation network for talking head video generation
Fa-Ting Hong and Dan Xu · 2023
Later among the works it cites.
Ae-nerf: Audio enhanced neural radiance field for few shot talking head synthesis
Dongze Li, Kang Zhao, Wei Wang, Bo Peng, Yingya Zhang, Jing Dong, and Tieniu Tan · 2023
Later among the works it cites.
Talkclip: Talking head generation with text-guided expressive speaking styles
Yifeng Ma, Suzhen Wang, Yu Ding, Bowen Ma, Tangjie Lv, Changjie Fan, Zhipeng Hu, Zhidong Deng, and Xin Yu · 2023
Later among the works it cites.
Styletalk: One-shot talking head generation with controllable speaking styles
Yifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding, Zhidong Deng, and Xin Yu · 2023
Later among the works it cites.
Emotional voice puppetry
Ye Pan, Ruisi Zhang, Shengran Cheng, Shuai Tan, Yu Ding, Kenny Mitchell, and Xubo Yang · 2023
Later among the works it cites.
Dpe: Disentanglement of pose and expression for general video portrait editing
Youxin Pang, Yong Zhang, Weize Quan, Yanbo Fan, Xiaodong Cun, Ying Shan, and Dong-ming Yan · 2023
Later among the works it cites.
Difftalk: Crafting diffusion models for generalized audio-driven portraits animation
Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu · 2023
Later among the works it cites.
Emmn: Emotional motion memory network for audio-driven emotional talking face generation
Shuai Tan, Bin Ji, and Ye Pan · 2023
Later among the works it cites.
Progressive disentangled representation learning for fine-grained controllable talking head synthesis
Duomin Wang, Yu Deng, Zixin Yin, Heung-Yeung Shum, and Baoyuan Wang · 2023
Later among the works it cites.
Seeing what you said: Talking face generation guided by a lip reading expert
Jiadong Wang, Xinyuan Qian, Malu Zhang, Robby T Tan, and Haizhou Li · 2023
Later among the works it cites.
Lipformer: High-fidelity and generalizable talking face generation with a pre-learned facial codebook
Jiayu Wang, Kang Zhao, Shiwei Zhang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou · 2023
Later among the works it cites.
Efficient video portrait reenactment via grid-based codebook
Kaisiyuan Wang, Hang Zhou, Qianyi Wu, Jiaxiang Tang, Zhiliang Xu, Borong Liang, Tianshu Hu, Errui Ding, Jingtuo Liu, Ziwei Liu, et al · 2023
Later among the works it cites.
Talking head generation with probabilistic audio-to-visual diffusion priors
Zhentao Yu, Zixin Yin, Deyu Zhou, Duomin Wang, Finn Wong, and Baoyuan Wang · 2023
Later among the works it cites.
Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation
Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang, Xi Shen, Yu Guo, Ying Shan, and Fei Wang · 2023
Later among the works it cites.
Identity-preserving talking face generation with landmark and appearance priors
Weizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei, Gangming Zhao, Liang Lin, and Guanbin Li · 2023
Later among the works it cites.
Expressive talking avatars
Ye Pan, Shuai Tan, Shengran Cheng, Qunfen Lin, Zijiao Zeng, and Kenny Mitchell · 2024
Closest in time.
Say anything with any style
Shuai Tan, Bin Ji, Yu Ding, and Ye Pan · 2024
Closest in time.
Shuai Tan, Bin Ji, and Ye Pan · 2024
Closest in time.
Style2talker: High-resolution talking head generation with emotion style and art style
Shuai Tan, Bin Ji, and Ye Pan · 2024
Closest in time.