Fetching the paper…
Reading the bibliography…
While recent research has made significant progress in speech-driven talking face generation, the quality of the generated video still lags behind that of real recordings.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. 2004 · 2004
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen. 2017 · 2017
Earlier work this paper cites.
Synthesizing Obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman. 2017 · 2017
Earlier work this paper cites.
A deep learning approach for generalized speech animation
Sarah Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews. 2017 · 2017
Earlier work this paper cites.
Lip Movements Generation at a Glance. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part VII (Lecture Notes in Computer Science, Vol. 11211) . 538–553
Lele Chen, Zhiheng Li, Ross K. Maddox, Zhiyao Duan, and Chenliang Xu. [n. d.] · 2018
Earlier work this paper cites.
The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In Proc. CVPR . 586–595
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. 2018 · 2018
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech. In Proc. ISCA Interspeech . 1526–1530
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019 · 2019
Earlier work this paper cites.
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations. In Proc. NeurIPS
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Earlier work this paper cites.
Denoising Diffusion Probabilistic Models. In Proc. NeurIPS
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM International Conference on Multimedia . 484–492
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. 2020 · 2020
Earlier work this paper cites.
Neural voice puppetry: Audio-driven facial reenactment. In Proc. ECCV . 716–731
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner. 2020 · 2020
Cited alongside, same era.
Realistic speech-driven facial animation with gans
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. 2020 · 2020
Cited alongside, same era.
Audio-driven talking face video generation with learning-based personalized head pose
Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao, and Yong-Jin Liu. 2020 · 2020
Cited alongside, same era.
Makelttalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. 2020 · 2020
Cited alongside, same era.
Diffusion Models Beat GANs on Image Synthesis. In Proc. NeurIPS . 8780–8794
Prafulla Dhariwal and Alexander Quinn Nichol. 2021 · 2021
Cited alongside, same era.
Latent Video Diffusion Models for High-Fidelity Video Generation with Arbitrary Lengths
Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. 2022 · 2022
Later among the works it cites.
Imagen video: High definition video generation with diffusion models
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al · 2022
Later among the works it cites.
Cascaded Diffusion Models for High Fidelity Image Generation
Jonathan Ho, Chitwan Saharia, William Chan, David J. Fleet, Mohammad Norouzi, and Tim Salimans. 2022b · 2022
Later among the works it cites.
Jonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. 2022c · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis. 5764–5774
Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang. 2021 · 2021
Cited alongside, same era.
LipSync3D: Data-Efficient Learning of Personalized 3D Talking Faces From Video Using Pose and Lighting Normalization. In Proc. CVPR . 2755–2764
Avisek Lahiri, Vivek Kwatra, Christian Früh, John Lewis, and Chris Bregler. 2021 · 2021
Cited alongside, same era.
Talking Head from Speech Audio using a Pre-trained Image Generator. In Proceedings of the 30th ACM International Conference on Multimedia . 5228–5236
Mohammed M Alghamdi, He Wang, Andrew J Bulpitt, and David C Hogg. 2022 · 2022
Cited alongside, same era.
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature. In Proc. ISCA Interspeech . 1596–1600
Chenpeng Du, Yiwei Guo, Xie Chen, and Kai Yu. 2022 · 2022
Cited alongside, same era.
Faceformer: Speech-driven 3d facial animation with transformers. In Proc. CVPR . 18770–18780
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura. 2022 · 2022
Cited alongside, same era.
SPACEx: Speech-driven Portrait Animation with Controllable Expression
Siddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle, and Ming-Yu Liu. 2022 · 2022
Cited alongside, same era.
Deep Unsupervised Learning using Nonequilibrium Thermodynamics. In Proc. ICML , Vol. 37. 2256–2265
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. [n. d.]
Cited in the paper.
Gyeongman Kim, Hajin Shim, Hyunsu Kim, Yunjey Choi, Junho Kim, and Eunho Yang. 2022 · 2022
Later among the works it cites.
Diffusion Autoencoders: Toward a Meaningful and Decodable Representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10619–10629
Konpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, and Supasorn Suwajanakorn. 2022 · 2022
Later among the works it cites.
Learning dynamic facial radiance fields for few-shot talking head synthesis. In Proc. ECCV . 666–682
Shuai Shen, Wanhua Li, Zheng Zhu, Yueqi Duan, Jie Zhou, and Jiwen Lu. 2022 · 2022
Later among the works it cites.
DFA-NERF: personalized talking head generation via disentangled face attributes neural rendering
Shunyu Yao, RuiZhe Zhong, Yichao Yan, Guangtao Zhai, and Xiaokang Yang. 2022 · 2022
Later among the works it cites.
Magicvideo: Efficient video generation with latent diffusion models
Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. 2022 · 2022
Later among the works it cites.
Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation
Michał Stypułkowski, Konstantinos Vougioukas, Sen He, Maciej Zięba, Stavros Petridis, and Maja Pantic. 2023 · 2023
Closest in time.