Fetching the paper…
Reading the bibliography…
Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved.
Bukimi no tani [the uncanny valley]
M. Mori · 1970
Earlier work this paper cites.
The DARPA speech recognition research database: Specifications and status
W. M. Fisher, G. R. Doddington, and K. M. Goudie-Marshall · 1986
Earlier work this paper cites.
Darpa timit acoustic phonetic continuous speech corpus cdrom, 1993
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, and N. L. Dahlgren · 1993
Earlier work this paper cites.
RASTA processing of speech
H. Hermansky and N. Morgan · 1994
Earlier work this paper cites.
Video rewrite: Driving visual speech with audio
C. Bregler, M. Covell, and M. Slaney · 1997
Earlier work this paper cites.
Voice puppetry
M. Brand · 1999
Earlier work this paper cites.
HMM-based text-to-audio-visual speech synthesis
S. Sako, K. Tokuda, T. Masuko, T. Kobayashi, and T. Kitamura · 2000
Earlier work this paper cites.
Speech driven facial animation
P. Kakumanu, R. Gutierrez-Osuna, A. Esposito, R. Bryll, A. Goshtasby, and O. Garcia · 2001
Earlier work this paper cites.
Trainable videorealistic speech animation
T. Ezzat, G. Geiger, and T. Poggio · 2002
Earlier work this paper cites.
Real-time speech-driven face animation with expressions using neural networks
P. Hong, Z. Wen, and T. S. Huang · 2002
Earlier work this paper cites.
Deformation transfer for triangle meshes
R.W. Sumner and J. Popović · 2004
Earlier work this paper cites.
Expressive speech-driven facial animation
Y. Cao, W. C. Tien, P. Faloutsos, and F. Pighin · 2005
Earlier work this paper cites.
Automatic 3D facial expression analysis in videos
Y. Chang, M. Vieira, M. Turk, and L. Velho · 2005
Earlier work this paper cites.
Face transfer with multilinear models
D. Vlasic, M. Brand, H. Pfister, and J. Popović · 2005
Earlier work this paper cites.
A 3D facial expression database for facial behavior research
L. Yin, X. Wei, Y. Sun, J. Wang, and M. J. Rosato · 2006
Earlier work this paper cites.
A 3D facial expression database for facial behavior research
L. Yin, X. Wei, Y. Sun, J. Wang, and M. J. Rosato · 2006
Earlier work this paper cites.
Rigid head motion in expressive speech animation: Analysis and synthesis
C. Busso, Z. Deng, M. Grimm, U. Neumann, and S. Narayanan · 2007
Earlier work this paper cites.
Realistic mouth-synching for speech-driven talking face using articulatory modelling
L. Xie and Z.-Q. Liu · 2007
Earlier work this paper cites.
Bosphorus database for 3D face analysis
A. Savran, N. Alyuöz, H. Dibeklioglu, O. Celiktutan, B. Gökberk, B. Sankur, and L. Akarun · 2008
Earlier work this paper cites.
A high-resolution 3D dynamic facial expression database
L. Yin, X. Chen, Y. Sun, T. Worm, and M. Reale · 2008
Earlier work this paper cites.
The digital emily project: Photoreal facial modeling and animation
O. Alexander, M. Rogers, W. Lambeth, M. Chiang, and P. Debevec · 2009
Earlier work this paper cites.
Synface—speech-driven facial animation for virtual speech-reading support
G. Salvi, J. Beskow, S. Al Moubayed, and B. Granström · 2009
Earlier work this paper cites.
A 3D audio-visual corpus of affective communication
Gabrielle Fanelli, Jürgen Gall, Harald Romsdorfer, Thibaut Weise, and Luc van Gool · 2010
Earlier work this paper cites.
Example-based facial rigging
H. Li, T. Weise, and M. Pauly · 2010
Cited alongside, same era.
A FACS valid 3D dynamic action unit database with applications to 3D dynamic morphable facial modeling
D. Cosker, E. Krumhuber, and A. Hilton · 2011
Cited alongside, same era.
Video face replacement
K. Dale, K. Sunkavalli, M. K. Johnson, D. Vlasic, W. Matusik, and H. Pfister · 2011
Cited alongside, same era.
Text driven 3d photo-realistic talking head
L. Wang, W. Han, F. Soong, and Q. Huo · 2011
Cited alongside, same era.
Realtime performance-based facial animation
T. Weise, S. Bouaziz, H. Li, and M. Pauly · 2011
Cited alongside, same era.
Dynamic units of visual speech
S. L. Taylor, M. Mahler, B.-J. Theobald, and I. Matthews · 2012
Cited alongside, same era.
JALI: An animator-centric viseme model for expressive lip synchronization
P. Edwards, C. Landreth, E. Fiume, and K. Singh · 2016
Later among the works it cites.
A deep bidirectional LSTM approach for video-realistic talking head
B. Fan, L. Xie, S. Yang, L. Wang, and F. K. Soong · 2016
Later among the works it cites.
Squad: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Later among the works it cites.
Audio-to-visual speech conversion using deep neural networks
S. Taylor, A. Kato, B. Milner, and I. Matthews · 2016
Later among the works it cites.
Face2Face: Real-time Face Capture and Reenactment of RGB Videos
J. Thies, M. Zollhöfer, M. Stamminger, C. Theobalt, and M. Nießner · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Facial expression editing in video using a temporally-smooth factorization
F. Yang, L. Bourdev, J. Wang, E. Shechtman, and D. Metaxas · 2012
Cited alongside, same era.
Expressive visual text-to-speech using active appearance models
R. Anderson, B. Stenger, V. Wan, and R. Cipolla · 2013
Cited alongside, same era.
Online modeling for realtime facial animation
S. Bouaziz, Y. Wang, and M. Pauly · 2013
Cited alongside, same era.
A new language independent, photo-realistic talking head driven by voice only
X. Zhang, L. Wang, G. Li, F. Seide, and F. K. Soong · 2013
Cited alongside, same era.
A 3D dynamic database for unconstrained face recognition
T. Alashkar, B. Ben Amor, M. Daoudi, and S. Berretti · 2014
Cited alongside, same era.
Displaced dynamic expression regression for real-time facial tracking and animation
C. Cao, Q. Hou, and K. Zhou · 2014
Cited alongside, same era.
A. van den Oord, S. Dieleman, H. Zen, . Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and . Kavukcuoglu · 2016
Later among the works it cites.
An anatomically-constrained local deformation model for monocular face capture
C. Wu, D. Bradley, M. Gross, and T. Beeler · 2016
Later among the works it cites.
Multimodal spontaneous emotion corpus for human behavior analysis
Z. Zhang, J. M. Girard, Y. Wu, X. Zhang, P. Liu, U. Ciftci, S. Canavan, M. Reale, A. Horowitz, H. Yang, J. F. Cohn, Q. Ji, and L. Yin · 2016
Later among the works it cites.
How far are we from solving the 2D & 3D face alignment problem? (and a dataset of 230,000 3D facial landmarks)
A. Bulat and G. Tzimiropoulos · 2017
Later among the works it cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen · 2017
Later among the works it cites.
Production-level facial performance capture using deep convolutional neural networks
S. Laine, T. Karras, T. Aila, A. Herva, S. Saito, R. Yu, H. Li, and J. Lehtinen · 2017
Later among the works it cites.
Learning a model of facial shape and expression from 4D scans
T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero · 2017
Later among the works it cites.
Project DeepSpeech
Mozilla · 2017
Later among the works it cites.
Speech-driven 3D facial animation with implicit emotional awareness: A deep learning approach
H. Pham and V. Pavlovic · 2017
Later among the works it cites.
Synthesizing Obama: learning lip sync from audio
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman · 2017
Later among the works it cites.
A deep learning approach for generalized speech animation
S. Taylor, T. Kim, Y. Yue, M. Mahler, J. Krahe, A. G. Rodriguez, J. Hodgins, and I. Matthews · 2017
Later among the works it cites.
Lip movements generation at a glance
L. Chen, Z. Li, R. K. Maddox, Z. Duan, and C. Xu · 2018
Later among the works it cites.
4DFAB: A large scale 4d database for facial expression analysis and biometric applications
S. Cheng, I. Kotsia, M. Pantic, and S. Zafeiriou · 2018
Later among the works it cites.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein · 2018
Later among the works it cites.
Generating 3D faces using convolutional mesh autoencoders
A. Ranjan, T. Bolkart, S. Sanyal, and M. J. Black · 2018
Later among the works it cites.
Visemenet: Audio-driven animator-centric speech animation
Y. Zhou, Y. Xu, C. Landreth, E. Kalogerakis, S. Maji, and K. Singh · 2018
Later among the works it cites.
Expressive body capture: 3d hands, face, and body from a single image
G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. A. Osman, D. Tzionas, and M. J. Black · 2019
Closest in time.
http://voca.is.tue.mpg.de , 2019
VOCA · 2019
Closest in time.