Fetching the paper…
Reading the bibliography…
Generating realistic talking faces is a complex and widely discussed task with numerous applications.
“Auto-encoding variational bayes,”
Diederik P Kingma and Max Welling, · 2013
Earlier work this paper cites.
“Deep speech 2: End-to-end speech recognition in english and mandarin,”
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al., · 2016
Earlier work this paper cites.
“Synthesizing obama: learning lip sync from audio,”
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman, · 2017
Earlier work this paper cites.
“Talking face generation by conditional recurrent adversarial network,”
Yang Song, Jingwen Zhu, Dawei Li, Xiaolong Wang, and Hairong Qi, · 2018
Earlier work this paper cites.
“Hierarchical cross-modal talking face generation with dynamic pixel-wise loss,”
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu, · 2019
Earlier work this paper cites.
“Capture, learning, and synthesis of 3d speaking styles,”
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black, · 2019
Earlier work this paper cites.
“A lip sync expert is all you need for speech to lip generation in the wild,”
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar, · 2020
Earlier work this paper cites.
“Makelttalk: speaker-aware talking-head animation,”
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li, · 2020
Cited alongside, same era.
“Lora: Low-rank adaptation of large language models,”
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen, · 2021
Cited alongside, same era.
“One-shot free-view neural talking-head synthesis for video conferencing,”
Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu, · 2021
Cited alongside, same era.
“Ad-nerf: Audio driven neural radiance fields for talking head synthesis,”
Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang, · 2021
Cited alongside, same era.
“Digital human resource development: where are we? where should we go and how do we go there?,”
Mohan Thite, · 2022
Cited alongside, same era.
“Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan,”
Fei Yin, Yong Zhang, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang, Qingyan Bai, Baoyuan Wu, Jue Wang, and Yujiu Yang, · 2022
Later among the works it cites.
“High-resolution image synthesis with latent diffusion models,”
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer, · 2022
Later among the works it cites.
“Stylesync: High-fidelity generalized and personalized lip sync in style-based generator,”
Jiazhi Guan, Zhanwang Zhang, Hang Zhou, Tianshu Hu, Kaisiyuan Wang, Dongliang He, Haocheng Feng, Jingtuo Liu, Errui Ding, Ziwei Liu, et al., · 2023
Closest in time.
“Adding conditional control to text-to-image diffusion models,”
Lvmin Zhang and Maneesh Agrawala, · 2023
Closest in time.
“Dae-talker: High fidelity speech-driven talking face generation with diffusion autoencoder,”
Chenpng Du, Qi Chen, Tianyu He, Xu Tan, Xie Chen, Kai Yu, Sheng Zhao, and Jiang Bian, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Eamm: One-shot emotional talking face via audio-based emotion-aware motion model,”
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Wayne Wu, Feng Xu, and Xun Cao, · 2022
Cited alongside, same era.
Closest in time.
“Audio-driven talking head video generation with diffusion model,”
Yizhe Zhua, Chunhui Zhanga, Qiong Liub, and Xi Zhoub, · 2023
Closest in time.