Fetching the paper…
Reading the bibliography…
We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli · 2004
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao · 2006
Earlier work this paper cites.
Modeling multimodal behaviors from speech prosody
Yu Ding, Catherine Pelachaud, and Thierry Artières · 2013
Earlier work this paper cites.
Lip reading in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Predicting head pose from speech with a conditional variational autoencoder
David Greenwood, Stephen Laycock, and Iain Matthews · 2017
Earlier work this paper cites.
Voxceleb: a large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman · 2017
Earlier work this paper cites.
Openface 2.0: Facial behavior analysis toolkit
Tadas Baltrusaitis, Amir Zadeh, Yao Chong Lim, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Recycle-gan: Unsupervised video retargeting
Aayush Bansal, Shugao Ma, Deva Ramanan, and Yaser Sheikh · 2018
Earlier work this paper cites.
Deep video portraits
Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Niessner, Patrick Pérez, Christian Richardt, Michael Zollhöfer, and Christian Theobalt · 2018
Earlier work this paper cites.
Novel realizations of speech-driven head movements with generative adversarial networks
Najmeh Sadoughi and Carlos Busso · 2018
Earlier work this paper cites.
Imagining the unimaginable faces by deconvolutional networks
Xin Yu and Fatih Porikli · 2018
Earlier work this paper cites.
Face super-resolution guided by facial component heatmaps
Xin Yu, Basura Fernando, Bernard Ghanem, Fatih Porikli, and Richard Hartley · 2018
Cited alongside, same era.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Cited alongside, same era.
Text-based editing of talking-head video
Ohad Fried, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shechtman, Dan B Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala · 2019
Cited alongside, same era.
Fsgan: Subject agnostic face swapping and reenactment
Yuval Nirkin, Yosi Keller, and Tal Hassner · 2019
Cited alongside, same era.
Animating arbitrary objects via deep motion transfer
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe · 2019
Cited alongside, same era.
First order motion model for image animation
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe · 2019
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Later among the works it cites.
Neural head reenactment with latent pose descriptors
Egor Burkov, Igor Pasechnik, Artur Grigorev, and Victor Lempitsky · 2020
Later among the works it cites.
Talking-head generation with rhythmic head motion
Lele Chen, Guofeng Cui, Celong Liu, Zhong Li, Ziyi Kou, Yi Xu, and Chenliang Xu · 2020
Later among the works it cites.
Marionette: Few-shot face reenactment preserving identity of unseen targets
Sungjoo Ha, Martin Kersner, Beomsu Kim, Seokjun Seo, and Dongyoung Kim · 2020
Later among the works it cites.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Later among the works it cites.
Neural voice puppetry: Audio-driven facial reenactment
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Talking face generation by conditional recurrent adversarial network
Yang Song, Jingwen Zhu, Dawei Li, Andy Wang, and Hairong Qi · 2019
Cited alongside, same era.
Realistic speech-driven facial animation with gans
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2019
Cited alongside, same era.
Few-shot video-to-video synthesis
Ting-Chun Wang, Ming-Yu Liu, Andrew Tao, Guilin Liu, Bryan Catanzaro, and Jan Kautz · 2019
Cited alongside, same era.
Semantic face hallucination: Super-resolving very low-resolution face images with supplementary attributes
Xin Yu, Basura Fernando, Richard Hartley, and Fatih Porikli · 2019
Cited alongside, same era.
Can we see more? joint frontalization and hallucination of unaligned tiny faces
Xin Yu, Fatemeh Shiri, Bernard Ghanem, and Fatih Porikli · 2019
Cited alongside, same era.
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner · 2020
Later among the works it cites.
Freenet: Multi-identity face reenactment
Jiangning Zhang, Xianfang Zeng, Mengmeng Wang, Yusu Pan, Liang Liu, Yong Liu, Yu Ding, and Changjie Fan · 2020
Later among the works it cites.
Makelttalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li · 2020
Later among the works it cites.
Arbitrary talking face generation via attentional audio-visual coherence learning
Hao Zhu, Huaibo Huang, Yi Li, Aihua Zheng, and Ran He · 2020
Later among the works it cites.
Write-a-speaker: Text-based emotional and rhythmic talking-head generation
Lincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding, Yixing Zheng, Xin Yu, and Changjie Fan · 2021
Closest in time.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan · 2021
Closest in time.