Fetching the paper…
Reading the bibliography…
Speech-driven 3D facial animation has been widely studied, yet there is still a gap to achieving realism and vividness due to the highly ill-posed nature and scarcity of audio-visual data.
Confusions among visually perceived consonants
Cletus G Fisher · 1968
Earlier work this paper cites.
Facial action coding system
Paul Ekman and Wallace V Friesen · 1978
Earlier work this paper cites.
Automated lip-sync: Background and techniques
John Lewis · 1991
Earlier work this paper cites.
Expressive speech-driven facial animation
Yong Cao, Wen C Tien, Petros Faloutsos, and Frédéric Pighin · 2005
Earlier work this paper cites.
Image denoising via learned dictionaries and sparse representation
Michael Elad and Michal Aharon · 2006
Earlier work this paper cites.
Computer facial animation
Frederic I Parke and Keith Waters · 2008
Earlier work this paper cites.
A 3-d audio-visual corpus of affective communication
Gabriele Fanelli, Juergen Gall, Harald Romsdorfer, Thibaut Weise, and Luc Van Gool · 2010
Earlier work this paper cites.
Realtime performance-based facial animation
Thibaut Weise, Sofien Bouaziz, Hao Li, and Mark Pauly · 2011
Earlier work this paper cites.
Animated speech: research progress and applications
DW Massaro, MM Cohen, M Tabain, J Beskow, and R Clark · 2012
Earlier work this paper cites.
Dynamic units of visual speech
Sarah L Taylor, Moshe Mahler, Barry-John Theobald, and Iain Matthews · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Realtime facial animation with on-the-fly correctives
Hao Li, Jihun Yu, Yuting Ye, and Chris Bregler · 2013
Earlier work this paper cites.
Anchored neighborhood regression for fast example-based super-resolution
Radu Timofte, Vincent De Smet, and Luc Van Gool · 2013
Earlier work this paper cites.
A practical and configurable lip sync method for games
Yuyu Xu, Andrew W Feng, Stacy Marsella, and Ari Shapiro · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
A+: Adjusted anchored neighborhood regression for fast super-resolution
Radu Timofte, Vincent De Smet, and Luc Van Gool · 2014
Earlier work this paper cites.
Convolutional sparse coding for image super-resolution
Shuhang Gu, Wangmeng Zuo, Qi Xie, Deyu Meng, Xiangchu Feng, and Lei Zhang · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Video-audio driven real-time facial animation
Yilong Liu, Feng Xu, Jinxiang Chai, Xin Tong, Lijuan Wang, and Qiang Huo · 2015
Earlier work this paper cites.
Audiovisual speech synthesis: An overview of the state-of-the-art
Wesley Mattheyses and Werner Verhelst · 2015
Earlier work this paper cites.
Real-time facial animation with image-based dynamic avatars
Chen Cao, Hongzhi Wu, Yanlin Weng, Tianjia Shao, and Kun Zhou · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Jali: an animator-centric viseme model for expressive lip synchronization
Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh · 2016
Earlier work this paper cites.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2016
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization
Xun Huang and Serge Belongie · 2017
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Cited alongside, same era.
Learning a model of facial shape and expression from 4d scans
Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero · 2017
Cited alongside, same era.
A deep learning approach for generalized speech animation
Sarah Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews · 2017
Cited alongside, same era.
Improved texture networks: Maximizing quality and diversity in feed-forward stylization and texture synthesis
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2017
Cited alongside, same era.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Audio-driven emotional video portraits
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu · 2021
Later among the works it cites.
Practical single-image super-resolution using look-up table
Younghyun Jo and Seon Joo Kim · 2021
Later among the works it cites.
Lipsync3d: Data-efficient learning of personalized 3d talking faces from video using pose and lighting normalization
Avisek Lahiri, Vivek Kwatra, Christian Frueh, John Lewis, and Chris Bregler · 2021
Later among the works it cites.
Generating diverse structure for image inpainting with hierarchical vq-vae
Jialun Peng, Dong Liu, Songcen Xu, and Houqiang Li · 2021
Later among the works it cites.
Audio-and gaze-driven facial animation of codec avatars
Alexander Richard, Colin Lea, Shugao Ma, Jurgen Gall, Fernando De la Torre, and Yaser Sheikh · 2021
Later among the works it cites.
Meshtalk: 3d face animation from speech using cross-modality disentanglement
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lip movements generation at a glance
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2018
Cited alongside, same era.
Deep video portraits
Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Niessner, Patrick Pérez, Christian Richardt, Michael Zollhöfer, and Christian Theobalt · 2018
Cited alongside, same era.
End-to-end learning for 3d facial animation from speech
Hai Xuan Pham, Yuting Wang, and Vladimir Pavlovic · 2018
Cited alongside, same era.
Visemenet: Audio-driven animator-centric speech animation
Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh · 2018
Cited alongside, same era.
State of the art on monocular 3d face reconstruction, tracking, and applications
Michael Zollhöfer, Justus Thies, Pablo Garrido, Derek Bradley, Thabo Beeler, Patrick Pérez, Marc Stamminger, Matthias Nießner, and Christian Theobalt · 2018
Cited alongside, same era.
Capture, learning, and synthesis of 3d speaking styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black · 2019
Cited alongside, same era.
Alexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre, and Yaser Sheikh · 2021
Later among the works it cites.
3d-talkemo: Learning to synthesize 3d emotional talking head
Qianyun Wang, Zhenfeng Fan, and Shihong Xia · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Later among the works it cites.
Talking head from speech audio using a pre-trained image generator
Mohammed M Alghamdi, He Wang, Andrew J Bulpitt, and David C Hogg · 2022
Later among the works it cites.
Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings
Tenglong Ao, Qingzhe Gao, Yuke Lou, Baoquan Chen, and Libin Liu · 2022
Later among the works it cites.
Faceformer: Speech-driven 3d facial animation with transformers
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura · 2022
Later among the works it cites.
Joint audio-text model for expressive speech-driven 3d facial animation
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura · 2022
Later among the works it cites.
Vector quantized diffusion model for text-to-image synthesis
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo · 2022
Later among the works it cites.
Unicolor: A unified framework for multi-modal colorization with transformer
Zhitong Huang, Nanxuan Zhao, and Jing Liao · 2022
Later among the works it cites.
Eamm: One-shot emotional talking face via audio-based emotion-aware motion model
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Wayne Wu, Feng Xu, and Xun Cao · 2022
Later among the works it cites.
Expressive talking head generation with granular audio-visual control
Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou, Zhibin Hong, Xiaoguang Han, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang · 2022
Later among the works it cites.
Semantic-aware implicit neural audio-driven video portrait generation
Xian Liu, Yinghao Xu, Qianyi Wu, Hang Zhou, Wayne Wu, and Bolei Zhou · 2022
Later among the works it cites.
Learning to listen: Modeling non-deterministic dyadic facial motion
Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li, Trevor Darrell, Angjoo Kanazawa, and Shiry Ginosar · 2022
Later among the works it cites.
Learning dynamic facial radiance fields for few-shot talking head synthesis
Shuai Shen, Wanhua Li, Zheng Zhu, Yueqi Duan, Jie Zhou, and Jiwen Lu · 2022
Later among the works it cites.
One-shot talking face generation from single-speaker audio-visual correlation learning
Suzhen Wang, Lincheng Li, Yu Ding, and Xin Yu · 2022
Later among the works it cites.
Audio-visual speech codecs: Rethinking audio-visual speech enhancement by re-synthesis
Karren Yang, Dejan Marković, Steven Krenn, Vasu Agrawal, and Alexander Richard · 2022
Later among the works it cites.
Metaportrait: Identity-preserving talking head generation with fast personalized adaptation
Bowen Zhang, Chenyang Qi, Pan Zhang, Bo Zhang, HsiangTao Wu, Dong Chen, Qifeng Chen, Yong Wang, and Fang Wen · 2022
Later among the works it cites.
Towards robust blind face restoration with codebook lookup transformer
Shangchen Zhou, Kelvin CK Chan, Chongyi Li, and Chen Change Loy · 2022
Later among the works it cites.
Dpe: Disentanglement of pose and expression for general video portrait editing
Youxin Pang, Yong Zhang, Weize Quan, Yanbo Fan, Xiaodong Cun, Ying Shan, and Dong-ming Yan · 2023
Closest in time.