Fetching the paper…
Reading the bibliography…
We present Neural Voice Puppetry, a novel approach for audio-driven facial video synthesis.
Bregler, C., Covell, M., Slaney, M.: Video rewrite: Driving visual speech with audio. In: Proceedings of the 24th Annual Conference on Computer Graphics and Interactive Techniques. pp. 353–360. SIGGRAPH ’97 (1997)
1997
Earlier work this paper cites.
Finn, K.E.: Video-Mediated Communication. L. Erlbaum Associates Inc. (1997)
1997
Earlier work this paper cites.
Blanz, V., Vetter, T.: A morphable model for the synthesis of 3D faces. In: ACM Transactions on Graphics (Proceedings of SIGGRAPH). pp. 187–194 (1999)
1999
Earlier work this paper cites.
Ezzat, T., Geiger, G., Poggio, T.: Trainable videorealistic speech animation. In: ACM Trans. on Graph. pp. 388–398 (2002)
2002
Earlier work this paper cites.
Tarasuik, J., Kaufman, J., Galligan, R.: Seeing is believing but is hearing? comparing audio and video communication for young children. Frontiers in Psychology 4
2013
Earlier work this paper cites.
Hannun, A., Case, C., Casper, J., Catanzaro, B., Diamos, G., Elsen, E., Prenger, R., Satheesh, S., Sengupta, S., Coates, A., Y. Ng, A.: DeepSpeech: Scaling up end-to-end speech recognition (12 2014)
2014
Earlier work this paper cites.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. CoRR abs/1412.6980
2014
Earlier work this paper cites.
Garrido, P., Valgaerts, L., Sarmadi, H., Steiner, I., Varanasi, K., Perez, P., Theobalt, C.: VDub - modifying face video of actors for plausible visual alignment to a dubbed audio track. In: Computer Graphics Forum (Proceedings of EUROGRAPHICS) (2015)
2015
Earlier work this paper cites.
Chung, J.S., Zisserman, A.: Out of time: automated lip sync in the wild. In: Workshop on Multi-view Lip-reading, ACCV (2016)
2016
Earlier work this paper cites.
Garrido, P., Zollhöfer, M., Casas, D., Valgaerts, L., Varanasi, K., Pérez, P., Theobalt, C.: Reconstruction of personalized 3D face rigs from monocular video. ACM Transactions on Graphics (Proceedings of SIGGRAPH) 35
2016
Earlier work this paper cites.
Isola, P., Zhu, J.Y., Zhou, T., Efros, A.A.: Image-to-image translation with conditional adversarial networks. arxiv (2016)
2016
Earlier work this paper cites.
Johnson, J., Alahi, A., Fei-Fei, L.: Perceptual losses for real-time style transfer and super-resolution. In: ECCV. pp. 694–711 (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Thies, J., Zollhöfer, M., Stamminger, M., Theobalt, C., Nießner, M.: Face2Face: Real-time Face Capture and Reenactment of RGB Videos. In: CVPR (2016)
2016
Cited alongside, same era.
Chung, J.S., Jamaludin, A., Zisserman, A.: You said that? In: British Machine Vision Conference (BMVC) (2017)
2017
Cited alongside, same era.
Karras, T., Aila, T., Laine, S., Herva, A., Lehtinen, J.: Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Trans. on Graph. (Proceedings of SIGGRAPH) 36
2017
Cited alongside, same era.
Li, T., Bolkart, T., Black, M.J., Li, H., Romero, J.: Learning a model of facial shape and expression from 4D scans. ACM Trans. on Graph. 36
2017
Cited alongside, same era.
Pham, H.X., Cheung, S., Pavlovic, V.: Speech-driven 3d facial animation with implicit emotional awareness: A deep learning approach. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2017, Honolulu, HI, USA, July 21-26, 2017. pp. 2328–2336. IEEE Computer Society (2017)
Thies, J., Zollhöfer, M., Stamminger, M., Theobalt, C., Nießner, M.: HeadOn: Real-time reenactment of human portrait videos. ACM Trans. on Graph. (Proceedings of SIGGRAPH) (2018)
2018
Later among the works it cites.
Vougioukas, K., Petridis, S., Pantic, M.: End-to-end speech-driven facial animation with temporal gans. In: BMVC (2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
Zollhöfer, M., Thies, J., Garrido, P., Bradley, D., Beeler, T., Pérez, P., Stamminger, M., Nießner, M., Theobalt, C.: State of the Art on Monocular 3D Face Reconstruction, Tracking, and Applications. Computer Graphics Forum (Eurographics State of the Art Reports) 37
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Suwajanakorn, S., Seitz, S.M., Kemelmacher-Shlizerman, I.: Synthesizing obama: Learning lip sync from audio. ACM Trans. on Graph. (Proceedings of SIGGRAPH) 36
2017
Cited alongside, same era.
Taylor, S., Kim, T., Yue, Y., Mahler, M., Krahe, J., Rodriguez, A.G., Hodgins, J., Matthews, I.: A deep learning approach for generalized speech animation. ACM Trans. on Graph. 36
2017
Cited alongside, same era.
Cozzolino, D., Thies, J., Rössler, A., Riess, C., Nießner, M., Verdoliva, L.: Forensictransfer: Weakly-supervised domain adaptation for forgery detection. arXiv (2018)
2018
Cited alongside, same era.
Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W.T., Rubinstein, M.: Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. ACM Trans. on Graph. 37
2018
Cited alongside, same era.
Jia, Y., Zhang, Y., Weiss, R.J., Wang, Q., Shen, J., Ren, F., Chen, Z., Nguyen, P., Pang, R., Moreno, I.L., Wu, Y.: Transfer learning from speaker verification to multispeaker text-to-speech synthesis. In: International Conference on Neural Information Processing Systems (NIPS). pp. 4485–4495 (2018)
2018
Cited alongside, same era.
Kim, H., Garrido, P., Tewari, A., Xu, W., Thies, J., Niessner, M., Pérez, P., Richardt, C., Zollhöfer, M., Theobalt, C.: Deep video portraits. ACM Trans. on Graph. (Proceedings of SIGGRAPH) 37
2018
Cited alongside, same era.
Lombardi, S., Saragih, J., Simon, T., Sheikh, Y.: Deep appearance models for face rendering. ACM Trans. on Graph. (Proceedings of SIGGRAPH) 37
2018
Cited alongside, same era.
Chen, L., Maddox, R.K., Duan, Z., Xu, C.: Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 7832–7841 (2019)
2019
Closest in time.
Cudeiro, D., Bolkart, T., Laidlaw, C., Ranjan, A., Black, M.: Capture, learning, and synthesis of 3D speaking styles. Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Closest in time.
Fried, O., Tewari, A., Zollhöfer, M., Finkelstein, A., Shechtman, E., Goldman, D.B., Genova, K., Jin, Z., Theobalt, C., Agrawala, M.: Text-based editing of talking-head video. ACM Trans. on Graph. (Proceedings of SIGGRAPH) 38
2019
Closest in time.
Kim, H., Elgharib, M., Zollhöfer, M., Seidel, H.P., Beeler, T., Richardt, C., Theobalt, C.: Neural style-preserving visual dubbing. ACM Trans. on Graph. (SIGGRAPH Asia) (2019)
2019
Closest in time.
Rössler, A., Cozzolino, D., Verdoliva, L., Riess, C., Thies, J., Nießner, M.: Faceforensics++: Learning to detect manipulated facial images. arXiv (2019)
2019
Closest in time.
Thies, J., Zollhöfer, M., Nießner, M.: Deferred neural rendering: Image synthesis using neural textures. ACM Trans. on Graph. (Proceedings of SIGGRAPH) (2019)
2019
Closest in time.
2019
Closest in time.
Vougioukas, K., Petridis, S., Pantic, M.: Realistic speech-driven facial animation with gans. International Journal of Computer Vision (IJCV) (Oct 2019)
2019
Closest in time.