Fetching the paper…
Reading the bibliography…
All previous methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone.
An audio-visual corpus for speech perception and automatic speech recognition
M. Cooke, J. Barker, S. Cunningham, and X. Shao · 2006
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, et al · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Photo-real talking head with deep bidirectional lstm
B. Fan, L. Wang, F. K. Soong, and L. Xie · 2015
Earlier work this paper cites.
Video-audio driven real-time facial animation
Y. Liu, F. Xu, J. Chai, X. Tong, L. Wang, and Q. Huo · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Lip reading in the wild
J. S. Chung and A. Zisserman · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild
J. S. Chung and A. Zisserman · 2016
Earlier work this paper cites.
Jali: an animator-centric viseme model for expressive lip synchronization
P. Edwards, C. Landreth, E. Fiume, and K. Singh · 2016
Earlier work this paper cites.
Jali: an animator-centric viseme model for expressive lip synchronization
P. Edwards, C. Landreth, E. Fiume, and K. Singh · 2016
Cited alongside, same era.
How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks)
A. Bulat and G. Tzimiropoulos · 2017
Cited alongside, same era.
You said that?
J. S. Chung, A. Jamaludin, and A. Zisserman · 2017
Cited alongside, same era.
Improved training of wasserstein gans
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville · 2017
Cited alongside, same era.
Unsupervised learning of disentangled and interpretable representations from sequential data
W.-N. Hsu, Y. Zhang, and J. Glass · 2017
Cited alongside, same era.
Image-to-image translation with conditional adversarial networks
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros · 2017
The conversation: Deep audio-visual speech enhancement
T. Afouras, J. S. Chung, and A. Zisserman · 2018
Later among the works it cites.
Lrs3-ted: a large-scale dataset for visual speech recognition
T. Afouras, J. S. Chung, and A. Zisserman · 2018
Later among the works it cites.
C. Chan, S. Ginosar, T. Zhou, and A. A. Efros · 2018
Later among the works it cites.
Lip movements generation at a glance
L. Chen, Z. Li, R. K Maddox, Z. Duan, and C. Xu · 2018
Later among the works it cites.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen · 2017
Cited alongside, same era.
Learning a model of facial shape and expression from 4D scans
T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero · 2017
Cited alongside, same era.
Synthesizing obama: learning lip sync from audio
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman · 2017
Cited alongside, same era.
A deep learning approach for generalized speech animation
S. Taylor, T. Kim, Y. Yue, M. Mahler, J. Krahe, A. G. Rodriguez, J. Hodgins, and I. Matthews · 2017
Cited alongside, same era.
Obamanet: Photo-realistic lip-sync from text
R. Kumar, J. Sotelo, K. Kumar, A. de Brébisson, and Y. Bengio
Cited in the paper.
S. A. Jalalifar, H. Hasani, and H. Aghajan · 2018
Later among the works it cites.
End-to-end speech-driven facial animation with temporal gans
K. Vougioukas, S. Petridis, and M. Pantic · 2018
Later among the works it cites.
Visemenet: Audio-driven animator-centric speech animation
Y. Zhou, Z. Xu, C. Landreth, E. Kalogerakis, S. Maji, and K. Singh · 2018
Later among the works it cites.
Capture, learning, and synthesis of 3D speaking styles
D. Cudeiro, T. Bolkart, C. Laidlaw, A. Ranjan, and M. Black · 2019
Closest in time.
Talking face generation by adversarially disentangled audio-visual representation
H. Zhou, Y. Liu, Z. Liu, P. Luo, and X. Wang · 2019
Closest in time.