Fetching the paper…
Reading the bibliography…
We present a novel approach to generating photo-realistic images of a face with accurate lip sync, given an audio input.
Live speech driven head-and-eye motion generators
Le, B.H., Ma, X., Deng, Z.: · 1914
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., Frasconi, P.: · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., Schmidhuber, J.: · 1997
Earlier work this paper cites.
Expressive speech-driven facial animation
Cao, Y., Tien, W.C., Faloutsos, P., Pighin, F.: · 2005
Earlier work this paper cites.
Audio signal feature extraction and classification using local discriminant bases
Umapathy, K., Krishnan, S., Rao, R.K.: · 2007
Earlier work this paper cites.
Unsupervised feature learning for audio classification using convolutional deep belief networks
Lee, H., Largman, Y., Pham, P., Ng, A.Y.: · 2009
Earlier work this paper cites.
Dlib-ml: A machine learning toolkit
King, D.E.: · 2009
Earlier work this paper cites.
An expressive text-driven 3d talking head
Anderson, R., Stenger, B., Wan, V., Cipolla, R.: · 2013
Earlier work this paper cites.
Supervised descent method and its applications to face alignment
Xiong, X., la Torre Frade, F.D.: · 2013
Earlier work this paper cites.
Automatic acquisition of high-fidelity facial performances using monocular videos
Shi, F., Wu, H.T., Tong, X., Chai, J.: · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
Mirza, M., Osindero, S.: · 2014
Earlier work this paper cites.
One millisecond face alignment with an ensemble of regression trees
Kazemi, V., Sullivan, J.: · 2014
Cited alongside, same era.
Extraction of features for lip-reading using autoencoders
Paleček, K.: · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D.P., Ba, J.: · 2014
Cited alongside, same era.
Talking heads synthesis from audio with deep neural networks
Shimba, T., Sakurai, R., Yamazoe, H., Lee, J.H.: · 2015
Cited alongside, same era.
Real-time expression transfer for facial reenactment
Thies, J., Zollhfer, M., Niessner, M., Valgaerts, L., Stamminger, M., Theobalt, C.: · 2015
Cited alongside, same era.
Vdub: Modifying face video of actors for plausible visual alignment to a dubbed audio track
Garrido, P., Valgaerts, L., Sarmadi, H., Steiner, I., Varanasi, K., Perez, P., Theobalt, C.: · 2015
Generating images with recurrent adversarial networks
Im, D.J., Kim, C.D., Jiang, H., Memisevic, R.: · 2016
Later among the works it cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., Abbeel, P.: · 2016
Later among the works it cites.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Zhang, H., Xu, T., Li, H., Zhang, S., Huang, X., Wang, X., Metaxas, D.N.: · 2016
Later among the works it cites.
Deconvolution and checkerboard artifacts
Odena, A., Dumoulin, V., Olah, C.: · 2016
Later among the works it cites.
Synthesizing obama: Learning lip sync from audio
Suwajanakorn, S., Seitz, S.M., Kemelmacher-Shlizerman, I.: · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Autoencoding beyond pixels using a learned similarity metric
Larsen, A.B.L., Sønderby, S.K., Winther, O.: · 2015
Cited alongside, same era.
Video-audio driven real-time facial animation
Liu, Y., Xu, F., Chai, J., Tong, X., Wang, L., Huo, Q.: · 2015
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems (2015) Software available from tensorflow.org
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G.S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., Zheng, X.: · 2015
Cited alongside, same era.
Web-based live speech-driven lip-sync
Llorach, G., Evans, A., Blat, J., Grimm, G., Hohmann, V.: · 2016
Cited alongside, same era.
Demo of face2face: Real-time face capture and reenactment of rgb videos
Thies, J., Zollhfer, M., Stamminger, M., Theobalt, C., Niessner, M.: · 2016
Cited alongside, same era.
Generative visual manipulation on the natural image manifold
Zhu, J.Y., Krähenbühl, P., Shechtman, E., Efros, A.A.: · 2016
Cited alongside, same era.
A deep learning approach for generalized speech animation
Taylor, S., Kim, T., Yue, Y., Mahler, M., Krahe, J., Rodriguez, A.G., Hodgins, J., Matthews, I.: · 2017
Later among the works it cites.
Progressive growing of gans for improved quality, stability, and variation
Karras, T., Aila, T., Laine, S., Lehtinen, J.: · 2017
Later among the works it cites.
Pose guided person image generation
Ma, L., Jia, X., Sun, Q., Schiele, B., Tuytelaars, T., Gool, L.V.: · 2017
Later among the works it cites.
Semi-latent GAN: learning to generate and modify facial images from attributes
Yin, W., Fu, Y., Sigal, L., Xue, X.: · 2017
Later among the works it cites.
Mocogan: Decomposing motion and content for video generation
Tulyakov, S., Liu, M., Yang, X., Kautz, J.: · 2017
Later among the works it cites.
Generative attribute controller with conditional filtered generative adversarial networks
Kaneko, T., Hiramatsu, K., Kashino, K.: · 2017
Later among the works it cites.
Aenet: Learning deep audio features for video analysis
Takahashi, N., Gygli, M., Gool, L.V.: · 2017
Later among the works it cites.