Fetching the paper…
Reading the bibliography…
We present a deep learning framework for real-time speech-driven 3D facial animation from just raw waveforms.
Long short-term memory
S. Hochreiter and J. Scmidhuber · 1997
Earlier work this paper cites.
Convolutional networks for images, speech, and time-series
Y. LeCun and Y. Bengio · 1998
Earlier work this paper cites.
Reanimating faces in images and video
V. Blanz, C. Basso, T. Poggio, and T. Vetter · 1999
Earlier work this paper cites.
Hmm-based text-to-audio-visual speech synthesis
S. Sako, K. Tokuda, T. Masuko, T. Kobayashi, and T. Kitamura · 2000
Earlier work this paper cites.
Trainable videorealistic speech animatio
T. Ezzat, G. Geiger, and T. Poggio · 2002
Earlier work this paper cites.
A morphable model for the synthesis of 3d faces
V. Blanz and T. Vetter · 2003
Earlier work this paper cites.
Lifelike talking faces for interactive services
E. Cosatto, J. Ostermann, H. P. Graf, and J. Schroeter · 2003
Earlier work this paper cites.
Expressive speech-driven facial animation
Y. Cao, W. C. Tien, P. Faloutsos, and F. Pighin · 2005
Earlier work this paper cites.
Real-time synthesis of chinese visual speech and facial expressions using mpeg-4 fap features in a three-dimensional avatar
Z. Wu, S. Zhang, L. Cai, and H. Meng · 2006
Earlier work this paper cites.
Video rewrite: driving visual speech with audio
C. Bregler, M. Covell, and M. Slaney · 2007
Earlier work this paper cites.
Assembling an expressive facial animation system
A. Wang, M. Emmi, and P. Faloutsos · 2007
Earlier work this paper cites.
Realistic mouth-synching for speech-driven talking face using articulatory modeling
L. Xie and Z. Liu · 2007
Earlier work this paper cites.
Synface: speech-driven facial animation for virtual speech-reading support
G. Salvi, J. Beskow, S. Moubayed, and B. Granstrom · 2009
Earlier work this paper cites.
Synthesizing photo-real talking head via trajectoryguided sample selection
L. Wang, X. Qian, W. Han, and F. K. Soong · 2010
Cited alongside, same era.
Text driven 3d photo-realistic talking head
L. Wang, X. Qian, F. K. Soong, and Q. Huo · 2011
Cited alongside, same era.
Applying convolutional neural networks concepts to hybrid nn-hmm model for speech recognition
O. Abdel-Hamid, A.-R. Mohamed, H. Jiang, and G. Penn · 2012
Cited alongside, same era.
Ravdess: The ryerson audio-visual database of emotional speech and song
S. R. Livingstone, K. Peck, and F. A. Russo · 2012
Cited alongside, same era.
A deep convolutional neural network using heterogeneous pooling for trading acoustic invariance with phonetic confusion
L. Deng, O. Abdel-Hamid, and D. Yu · 2013
Cited alongside, same era.
Estimating phoneme class conditional probabilities from raw speech signal using convolutional neural networks
Speech acoustic modeling from raw multichannel waveforms
Y. Hoshen, R. J. Weiss, and K. W. Wilson · 2015
Later among the works it cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Later among the works it cites.
Deep convolutional neural networks for large-scale speech tasks
T. N. Sainath, B. Kingsbury, G. Saon, H. Soltau, A. rahman Mohamed, G. Dahl, and B. Ramabhadran · 2015
Later among the works it cites.
Convolutional, long short-term memory, fully connected deep neural networks
T. N. Sainath, O. Vinyals, A. Senior, and H. Sak · 2015
Later among the works it cites.
Learning the speech front-end with raw waveforms cldnns
T. N. Sainath, R. J. Weiss, A. Senior, K. W. Wilson, and O. Vinyals · 2015
Later among the works it cites.
A deep bidirectional lstm approach for video-realistic talking head
B. Fan, L. Xie, S. Yang, L. Wang, and F. K. Soong · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Palaz, R. Collobert, and M. Magimai-Doss · 2013
Cited alongside, same era.
Statistical parametric speech synthesis using deep neural networks
H. Zen, A. Senior, and M. Schuster · 2013
Cited alongside, same era.
A new language independent, photo realistic talking head driven by voice only
X. Zhang, L. Wang, G. Li, F. Seide, and F. K. Soong · 2013
Cited alongside, same era.
Convolutional neural networks for speech recognition
O. Abdel-Hamid, A.-R. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu · 2014
Cited alongside, same era.
FaceWarehouse: A 3D Facial Expression Database for Visual Computing
C. Cao, Y. Weng, S. Zhou, Y. Tong, and K. Zhou · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
On the training aspects of deep neural network (dnn) for parametric tts synthesis
Y. Qian, Y. Fan, and F. K. Soong · 2014
Cited alongside, same era.
Later among the works it cites.
Robust real-time 3d face tracking from rgbd videos under extreme pose, depth, and expression variations
H. X. Pham and V. Pavlovic · 2016
Later among the works it cites.
Robust real-time performance-driven 3d face tracking
H. X. Pham, V. Pavlovic, J. Cai, and T. jen Cham · 2016
Later among the works it cites.
Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network
G. Trigeorgis, F. Ringeval, R. Brueckner, E. Marchi, M. A. Nicolaou, B. Schuller, and S. Zafeiriou · 2016
Later among the works it cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen · 2017
Closest in time.
Speech-driven 3d facial animation with implicit emotional awareness: a deep learning approach
H. X. Pham, S. Cheung, and V. Pavlovic · 2017
Closest in time.
Synthesizing obama: learning lip sync from audio
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Schlizerman · 2017
Closest in time.
A deep learning approach for generalized speech animation
S. Taylor, T. Kim, Y. Yue, M. Mahler, J. Krahe, A. G. Rodriguez, J. Hodgins, and I. Matthews · 2017
Closest in time.