Fetching the paper…
Reading the bibliography…
Given an arbitrary face image and an arbitrary speech clip, the proposed work attempts to generating the talking face video with accurate lip synchronization while maintaining smooth transition of both lip and facial movement over the entire video clip.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Computer vision aided lip movement correction to improve english pronunciation
Shuang Wei · 2014
Earlier work this paper cites.
Tcd-timit: An audio-visual corpus of continuous speech
Naomi Harte and Eoin Gillen · 2015
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Lip reading in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
You said that?
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman · 2017
Cited alongside, same era.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Cited alongside, same era.
Dual motion gan for future-flow embedded video prediction
Xiaodan Liang, Lisa Lee, Wei Dai, and Eric P Xing · 2017
Cited alongside, same era.
Voxceleb: a large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman · 2017
Cited alongside, same era.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman · 2017
Mocogan: Decomposing motion and content for video generation
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz · 2017
Later among the works it cites.
Decomposing motion and content for natural video sequence prediction
Ruben Villegas, Jimei Yang, Seunghoon Hong, Xunyu Lin, and Honglak Lee · 2017
Later among the works it cites.
Age progression/regression by conditional adversarial autoencoder
Zhifei Zhang, Yang Song, and Hairong Qi · 2017
Later among the works it cites.
Lip movements generation at a glance
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2018
Closest in time.
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A deep learning approach for generalized speech animation
Sarah Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews · 2017
Cited alongside, same era.
Closest in time.
Image super-resolution by neural texture transfer
Zhifei Zhang, Zhaowen Wang, Zhe Lin, and Hairong Qi · 2019
Closest in time.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Closest in time.