Fetching the paper…
Reading the bibliography…
The task of talking head generation is to synthesize a lip synchronized talking head video by inputting an arbitrary face image and audio clips.
Facial action coding system: A technique for the measurement of facial movement
P. Ekman and W. Friesen · 1978
Earlier work this paper cites.
Video rewrite: Driving visual speech with audio
Christoph Bregler, Michele Covell, and Malcolm Slaney · 1997
Earlier work this paper cites.
Lip movement synthesis from speech based on hidden markov models
Eli Yamamoto, Satoshi Nakamura, and Kiyohiro Shikano · 1998
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli · 2004
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao · 2006
Earlier work this paper cites.
A coupled HMM approach to video-realistic speech animation
Lei Xie and Zhiqiang Liu · 2007
Earlier work this paper cites.
Dlib-ml: A machine learning toolkit
Davis E King · 2009
Earlier work this paper cites.
Generative adversarial networks
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Cross-dataset learning and person-specific normalisation for automatic action unit detection
Tadas Baltrušaitis, Marwa Mahmoud, and Peter Robinson · 2015
Earlier work this paper cites.
Photo-real talking head with deep bidirectional lstm
Bo Fan, Lijuan Wang, Frank K Soong, and Lei Xie · 2015
Earlier work this paper cites.
TCD-TIMIT: An audio-visual corpus of continuous speech
Naomi Harte and Eoin Gillen · 2015
Earlier work this paper cites.
Face reading from speech—predicting facial action units from audio cues
Fabien Ringeval, Erik Marchi, Marc Mehu, Klaus Scherer, and Björn Schuller · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Cited alongside, same era.
Lip reading in the wild
Joon Son Chung and Andrew Zisserman · 2016
Cited alongside, same era.
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei · 2016
Cited alongside, same era.
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Fausto Milletari, Nassir Navab, and Seyed Ahmad Ahmadi · 2016
Cited alongside, same era.
You said that?
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman · 2017
Cited alongside, same era.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros · 2017
Cited alongside, same era.
End-to-end speech-driven facial animation with temporal gans
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2018
Later among the works it cites.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Later among the works it cites.
You said that?: Synthesising talking faces from audio
Amir Jamaludin, Joon Son Chung, and Andrew Zisserman · 2019
Later among the works it cites.
Face inpainting with dynamic structural information of facial action units
Le Li, Zhilei Liu, and Cuicui Zhang · 2019
Later among the works it cites.
Talking face generation by conditional recurrent adversarial network
Yang Song, Jingwen Zhu, Dawei Li, Andy Wang, and Hairong Qi · 2019
Later among the works it cites.
Realistic speech-driven facial animation with GANs
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automatic analysis of facial actions: A survey
Brais Martinez, Michel F Valstar, Bihan Jiang, and Maja Pantic · 2017
Cited alongside, same era.
Listen to your face: Inferring facial action units from audio channel
Zibo Meng, Shizhong Han, and Yan Tong · 2017
Cited alongside, same era.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman · 2017
Cited alongside, same era.
Photorealistic facial expression synthesis by the conditional difference adversarial autoencoder
Yuqian Zhou and Bertram Emil Shi · 2017
Cited alongside, same era.
Openface 2.0: Facial behavior analysis toolkit
Tadas Baltrušaitis, Amir Zadeh, Yao Chong Lim, and Louis-Philippe Morency · 2018
Cited alongside, same era.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Cited alongside, same era.
Later among the works it cites.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Later among the works it cites.
End-to-end generation of talking faces from noisy speech
Sefik Emre Eskimez, Ross K Maddox, Chenliang Xu, and Zhiyao Duan · 2020
Later among the works it cites.
Region based adversarial synthesis of facial action units
Zhilei Liu, Diyi Liu, and Yunpeng Wu · 2020
Later among the works it cites.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Later among the works it cites.
Makelttalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li · 2020
Later among the works it cites.
Speech driven talking face generation from a single image and an emotion condition
Sefik Emre Eskimez, You Zhang, and Zhiyao Duan · 2021
Closest in time.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Closest in time.