Fetching the paper…
Reading the bibliography…
Speech-driven facial animation is the process that automatically synthesizes talking characters based on speech signals.
Proceedings of the Institute of Acoustics, Autumn Meeting 12
Simons, A.D., Cox, S.J.: Generation of mouthshapes for a synthetic talking head · 1990
Earlier work this paper cites.
Movement Disorders 12
Bentivoglio, A.R., Bressman, S.B., Cassetta, E., Carretta, D., Tonali, P., Albanese, A.: Analysis of blink rate patterns in normal subjects · 1997
Earlier work this paper cites.
In: Proceedings of the 24th Annual Conference on Computer Graphics and Interactive Techniques, pp. 353–360 (1997)
Bregler, C., Covell, M., Slaney, M.: Video Rewrite · 1997
Earlier work this paper cites.
Speech Communication 26
Yamamoto, E., Nakamura, S., Shikano, K.: Lip movement synthesis from speech based on hidden Markov Models · 1998
Earlier work this paper cites.
Speech Communication 26
Yehia, H., Rubin, P., Vatikiotis-Bateson, E.: Quantitative association of vocal-tract and facial behavior · 1998
Earlier work this paper cites.
Journal of Phonetics 30
Yehia, H.C., Kuratate, T., Vatikiotis-Bateson, E.: Linking facial animation, head motion and speech acoustics · 2002
Earlier work this paper cites.
ACM TOG 24
Cao, Y., Tien, W.C., Faloutsos, P., Pighin, F.: Expressive speech-driven facial animation · 2005
Earlier work this paper cites.
The Journal of the Acoustical Society of America 120
Cooke, M., Barker, J., Cunningham, S., Shao, X.: An audio-visual corpus for speech perception and automatic speech recognition · 2006
Earlier work this paper cites.
Pattern Recognition 40
Xie, L., Liu, Z.Q.: A coupled HMM approach to video-realistic speech animation · 2007
Earlier work this paper cites.
International Workshop on Quality of Multimedia Experience (QoMEx) 20
Narvekar, N.D., Karam, L.J.: A no-reference perceptual image sharpness metric based on a cumulative probability of blur detection · 2009
Earlier work this paper cites.
IEEE Transactions on Affective Computing 5
Cao, H., Cooper, D.G., Keutmann, M.K., Gur, R.C., Nenkova, A., Verma, R.: CREMA-D: Crowd-sourced emotional multimodal actors dataset · 2014
Earlier work this paper cites.
In: NIPS, pp. 2672–2680 (2014)
Goodfellow, I.J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative Adversarial Networks · 2014
Earlier work this paper cites.
arXiv preprint arXiv:1412.6980 (2014)
Kingma, D.P., Ba, J.: Adam: A Method for Stochastic Optimization · 2014
Earlier work this paper cites.
In: ICASSP, pp. 4884–4888 (2015)
Fan, B., Wang, L., Soong, F., Xie, L.: Photo-real talking head with deep bidirectional lstm · 2015
Earlier work this paper cites.
IEEE Transactions on Multimedia 17
Harte, N., Gillen, E.: TCD-TIMIT: An audio-visual corpus of continuous speech · 2015
Earlier work this paper cites.
arXiv preprint arXiv:1511.05440 (2015)
Mathieu, M., Couprie, C., LeCun, Y.: Deep multi-scale video prediction beyond mean square error · 2015
Cited alongside, same era.
arXiv preprint arXiv:1511.06434 (2015)
Radford, A., Metz, L., Chintala, S.: Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks · 2015
Cited alongside, same era.
In: International Conference on Medical image computing and computer-assisted intervention, pp. 234–241 (2015)
Ronneberger, O., Fischer, P., Brox, T.: U-Net: Convolutional Networks for Biomedical Image Segmentation · 2015
Cited alongside, same era.
Tech. Rep. 118 (2016)
Amos, B., Ludwiczuk, B., Satyanarayanan, M.: OpenFace: A general-purpose face recognition library with mobile applications · 2016
Cited alongside, same era.
In: CVPR-Workshop, pp. 2328–2336 (2017)
Pham, H.X., Cheung, S., Pavlovic, V.: Speech-Driven 3D Facial Animation with Implicit Emotional Awareness: A Deep Learning Approach · 2017
Later among the works it cites.
In: ICCV, pp. 2830–2839 (2017)
Saito, M., Matsumoto, E., Saito, S.: Temporal Generative Adversarial Nets with Singular Value Clipping · 2017
Later among the works it cites.
ACM TOG 36
Suwajanakorn, S., Seitz, S., Kemelmacher-Shlizerman, I.: Synthesizing Obama: Learning Lip Sync from Audio Output Obama Video · 2017
Later among the works it cites.
ACM TOG 36
Taylor, S., Kim, T., Yue, Y., Mahler, M., Krahe, J., Rodriguez, A.G., Hodgins, J., Matthews, I.: A deep learning approach for generalized speech animation · 2017
Later among the works it cites.
arXiv preprint arXiv:1707.04993 (2017)
Tulyakov, S., Liu, M., Yang, X., Kautz, J.: MoCoGAN: Decomposing Motion and Content for Video Generation · 2017
Later among the works it cites.
IEEE TPAMI (2017)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Assael, Y.M., Shillingford, B., Whiteson, S., de Freitas, N.: LipNet: End-to-End Sentence-level Lipreading · 2016
Cited alongside, same era.
In: ACCV (2016)
Chung, J.S., Zisserman, A.: Lip reading in the wild · 2016
Cited alongside, same era.
In: Workshop on Multi-view Lip-reading, ACCV (2016)
Chung, J.S., Zisserman, A.: Out of time: automated lip sync in the wild · 2016
Cited alongside, same era.
In: Computer Vision Winter Workshop (2016)
Soukupova, T., Cech, J.: Real-time eye blink detection using facial landmarks · 2016
Cited alongside, same era.
In: NIPS, pp. 613–621 (2016)
Vondrick, C., Pirsiavash, H., Torralba, A.: Generating Videos with Scene Dynamics · 2016
Cited alongside, same era.
In: ICLR (2017)
Arjovsky, M., Bottou, L.: Towards Principled Methods for Training Generative Adversarial Networks · 2017
Cited alongside, same era.
In: Thematic Workshops of ACM Multimedia, pp. 349–357 (2017)
Chen, L., Srivastava, S., Duan, Z., Xu, C.: Deep Cross-Modal Audio-Visual Generation · 2017
Cited alongside, same era.
In: BMVC (2017)
Chung, J.S., Jamaludin, A., Zisserman, A.: You said that? · 2017
Cited alongside, same era.
Zhu, X., Lei, Z., Li, S.Z., et al.: Face alignment in full pose range: A 3d total solution · 2017
Later among the works it cites.
In: ECCV, pp. 1–15 (2018)
Chen, L., Li, Z., Maddox, R.K., Duan, Z., Xu, C.: Lip Movements Generation at a Glance · 2018
Later among the works it cites.
https://github.com/cleardusk/3DDFA
Jianzhu Guo, X.Z., Lei, Z.: 3ddfa · 2018
Later among the works it cites.
arXiv preprint arXiv:1806.02877 (2018)
Li, Y., Chang, M., Lyu, S.: In ictu oculi : Exposing ai generated fake face videos by detecting eye blinking · 2018
Later among the works it cites.
Pham, H.X., Wang, Y., Pavlovic, V.: Generative Adversarial Talking Head: Bringing Portraits to Life with a Weakly Supervised Neural Network pp. 1–18 (2018)
2018
Later among the works it cites.
In: ECCV (2018)
Pumarola, A., Agudo, A., Martinez, A., Sanfeliu, A., Moreno-Noguer, F.: Ganimation: Anatomically-aware facial animation from a single image · 2018
Later among the works it cites.
In: BMVC (2018)
Vougioukas, K., Petridis, S., Pantic, M.: End-to-End Speech-Driven Facial Animation with Temporal GANs · 2018
Later among the works it cites.
ACM TOG 37
Zhou, Y., Xu, Z., Landreth, C., Kalogerakis, E., Maji, S., Singh, K.: VisemeNet: Audio-Driven Animator-Centric Speech Animation · 2018
Later among the works it cites.
In: CVPR (2019)
Lele Chen Ross K Maddox, Z.D.C.X.: Hierarchical cross-modal talking face generation with dynamic pixel-wise loss · 2019
Closest in time.
In: AAAI (2019)
Zhou, H., Liu, Y., Liu, Z., Luo, P., Wang, X.: Talking face generation by adversarially disentangled audio-visual representation · 2019
Closest in time.