Fetching the paper…
Reading the bibliography…
This paper presents a generic method for generating full facial 3D animation from speech.
Video rewrite: Driving visual speech with audio
Christoph Bregler, Michele Covell, and Malcolm Slaney · 1997
Earlier work this paper cites.
Miketalk: A talking facial display based on morphing visemes
Tony Ezzat and Tomaso Poggio · 1998
Earlier work this paper cites.
Voice puppetry
Matthew Brand · 1999
Earlier work this paper cites.
Voice puppetry
Matthew Brand · 1999
Earlier work this paper cites.
Visual speech synthesis by morphing visemes
Tony Ezzat and Tomaso Poggio · 2000
Earlier work this paper cites.
Face animation based on observed 3d speech dynamics
Gregor A Kalberer and Luc Van Gool · 2001
Earlier work this paper cites.
Speech animation using viseme space
Gregor A Kalberer, Pascal Müller, and Luc Van Gool · 2002
Earlier work this paper cites.
Using viseme based acoustic models for speech driven lip synthesis
Ashish Verma, Nitendra Rajput, and L Venkata Subramaniam · 2003
Earlier work this paper cites.
Expressive speech-driven facial animation
Yong Cao, Wen C Tien, Petros Faloutsos, and Frédéric Pighin · 2005
Earlier work this paper cites.
Facial animation based on context-dependent visemes
José Mario De Martino, Léo Pini Magalhães, and Fábio Violaro · 2006
Earlier work this paper cites.
Animating blendshape faces by cross-mapping motion capture data
Zhigang Deng, Pei-Ying Chiang, Pamela Fox, and Ulrich Neumann · 2006
Earlier work this paper cites.
High quality lip-sync animation for 3d photo-realistic talking head
Lijuan Wang, Wei Han, and Frank K Soong · 2012
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al · 2014
Earlier work this paper cites.
JALI: An animator-centric viseme model for expressive lip synchronization
Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh · 2016
Earlier work this paper cites.
Reconstruction of personalized 3d face rigs from monocular video
Pablo Garrido, Michael Zollhöfer, Dan Casas, Levi Valgaerts, Kiran Varanasi, Patrick Pérez, and Christian Theobalt · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aäron Van Den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Conditional image generation with pixelcnn decoders
Aaron van den Oord, Nal Kalchbrenner, Lasse Espeholt, koray kavukcuoglu, Oriol Vinyals, and Alex Graves · 2016
Cited alongside, same era.
You said that?
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman · 2017
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Cited alongside, same era.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
End-to-end learning for 3d facial animation from speech
Hai Xuan Pham, Yuting Wang, and Vladimir Pavlovic · 2018
Later among the works it cites.
End-to-end speech-driven facial animation with temporal gans
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2018
Later among the works it cites.
X2face: A network for controlling face generation using images, audio, and pose codes
Olivia Wiles, A Koepke, and Andrew Zisserman · 2018
Later among the works it cites.
Visemenet: Audio-driven animator-centric speech animation
Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh · 2018
Later among the works it cites.
Capture, learning, and synthesis of 3d speaking styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J. Black · 2019
Later among the works it cites.
Variational mixture-of-experts autoencoders for multi-modal deep generative models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Synthesizing obama: Learning lip sync from audio
Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman · 2017
Cited alongside, same era.
A deep learning approach for generalized speech animation
Sarah Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
https://aws.amazon.com/sumerian/
Amazon sumerian, 2018 · 2018
Cited alongside, same era.
Generating talking face landmarks from speech
Sefik Emre Eskimez, Ross K. Maddox, Chenliang Xu, and Zhiyao Duan · 2018
Cited alongside, same era.
Joint learning of facial expression and head pose from speech
David Greenwook, Iain Matthews, and Stephen D. Laycock · 2018
Cited alongside, same era.
Yuge Shi, Narayanaswamy Siddharth, Brooks Paige, and Philip Torr · 2019
Later among the works it cites.
Talking face generation by conditional recurrent adversarial network
Yang Song, Jingwen Zhu, Xiaolong Wang, and Hairong Qi · 2019
Later among the works it cites.
VR facial animation via multiview image translation
Shih-En Wei, jason Saragih, Tomas Simon, Adam W. Harley, Stephen Lombardi, Michal Purdoch, Alexander Hypes, Dawei Wang, Hernan Badino, and Yaser Sheikh · 2019
Later among the works it cites.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Later among the works it cites.
Expressive telepresence via modular codec avatars
Hang Chu, Shugao Ma, Fernando De la Torre, Sanja Fidler, and Yaser Sheikh · 2020
Later among the works it cites.
A lip sync expert is all you need for speech to lip generation in the wild
K R Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and C V Jawahar · 2020
Later among the works it cites.
Neural voice puppetry: Audio-driven facial reenactment
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner · 2020
Later among the works it cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Later among the works it cites.
Audio- and gaze-driven facial animation of codec avatars
Alexander Richard, Colin Lea, Shugao Ma, Juergen Gall, Fernando de la Torre, and Yaser Sheikh · 2021
Closest in time.