Fetching the paper…
Reading the bibliography…
Generating 3D speech-driven talking head has received more and more attention in recent years.
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for HMM-based speech synthesis,” in
2000
Earlier work this paper cites.
D. Schabus, M. Pucher, and G. Hofer, “Simultaneous speech and animation synthesis,” in
2011
Earlier work this paper cites.
B.-J. Theobald and I. Matthews, “Relating objective and subjective performance measures for aam-based visual speech synthesis,”
2012
Earlier work this paper cites.
S. L. Taylor, M. Mahler, B.-J. Theobald, and I. Matthews, “Dynamic units of visual speech,” in
2012
Earlier work this paper cites.
L. Wang, W. Han, and F. K. Soong, “High quality lip-sync animation for 3d photo-realistic talking head,” in
2012
Earlier work this paper cites.
W. Mattheyses, L. Latacz, and W. Verhelst, “Comprehensive many-to-many phoneme-to-viseme mapping and its application for concatenative visual speech synthesis,”
2013
Earlier work this paper cites.
Y. Xu, A. W. Feng, S. Marsella, and A. Shapiro, “A practical and configurable lip sync method for games,” in
2013
Earlier work this paper cites.
T. Merritt and S. King, “Investigating the shortcomings of HMM synthesis,” in
2013
Earlier work this paper cites.
C. Cao, Y. Weng, S. Zhou, Y. Tong, and K. Zhou, “Facewarehouse: A 3D facial expression database for visual computing,”
2013
Earlier work this paper cites.
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman, “Return of the devil in the details: Delving deep into convolutional nets,” in
2014
Cited alongside, same era.
P. Edwards, C. Landreth, E. Fiume, and K. Singh, “Jali: an animator-centric viseme model for expressive lip synchronization,”
2016
Cited alongside, same era.
B. Fan, L. Xie, S. Yang, L. Wang, and F. K. Soong, “A deep bidirectional LSTM approach for video-realistic talking head,”
2016
Cited alongside, same era.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,” in
2016
Cited alongside, same era.
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Synthesizing obama: learning lip sync from audio,”
2017
Cited alongside, same era.
2018
Later among the works it cites.
D. Greenwood, I. Matthews, and S. Laycock, “Joint learning of facial expression and head pose from speech,” in
2018
Later among the works it cites.
H. Kim, P. Garrido, A. Tewari, W. Xu, J. Thies, M. Nießner, P. Pérez, C. Richardt, M. Zollhöfer, and C. Theobalt, “Deep video portraits,”
2018
Later among the works it cites.
F. Wu, L. Bao, Y. Chen, Y. Ling, Y. Song, S. Li, K. N. Ngan, and W. Liu, “Mvf-net: Multi-view 3d face morphable model regression,” in
2019
Later among the works it cites.
T. Biasutto, S. Dahmani, S. Ouni
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Taylor, T. Kim, Y. Yue, M. Mahler, J. Krahe, A. G. Rodriguez, J. Hodgins, and I. Matthews, “A deep learning approach for generalized speech animation,”
2017
Cited alongside, same era.
J. S. Chung, A. Jamaludin, and A. Zisserman, “You said that?” in
2017
Cited alongside, same era.
D. Krueger, T. Maharaj, J. Kramár, M. Pezeshki, N. Ballas, N. R. Ke, A. Goyal, Y. Bengio, A. C. Courville, and C. J. Pal, “Zoneout: Regularizing rnns by randomly preserving hidden activations,” in
2017
Cited alongside, same era.
S. Dahmani, V. Colotte, V. Girard, and S. Ouni, “Conditional variational auto-encoder for text-driven expressive audiovisual speech synthesis,” in
2019
Later among the works it cites.
H. Lu, Z. Wu, R. Li, S. Kang, J. Jia, and H. Meng, “A compact framework for voice conversion using wavenet conditioned on phonetic posteriorgrams,” in
2019
Later among the works it cites.
O. Fried, A. Tewari, M. Zollhöfer, A. Finkelstein, E. Shechtman, D. B. Goldman, K. Genova, Z. Jin, C. Theobalt, and M. Agrawala, “Text-based editing of talking-head video,”
2019
Later among the works it cites.