Fetching the paper…
Reading the bibliography…
Speech-driven 3D facial animation plays a key role in applications such as virtual avatars, gaming, and digital content creation.
Capture, Learning, and Synthesis of 3D Speaking Styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J. Black. 2019 · 1905
Earlier work this paper cites.
MediaPipe: A Framework for Building Perception Pipelines
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, Wan-Teh Chang, Wei Hua, Manfred Georg, and Matthias Grundmann. 2019 · 1906
Earlier work this paper cites.
Confusions Among Visually Perceived Consonants
Cletus G. Fisher. 1968 · 1968
Earlier work this paper cites.
Modeling Coarticulation in Synthetic Visual Speech. In Models and Techniques in Computer Animation , Nadia Magnenat Thalmann and Daniel Thalmann (Eds.). Springer Japan, Tokyo, 139–156
Michael M. Cohen and Dominic W. Massaro. 1993 · 1993
Earlier work this paper cites.
MakeItTalk: Speaker-Aware Talking-Head Animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. 2020 · 2004
Earlier work this paper cites.
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2006
Earlier work this paper cites.
A 3-D Audio-Visual Corpus of Affective Communication
Gabriele Fanelli, Juergen Gall, Harald Romsdorfer, Thibaut Weise, and Luc Van Gool. 2010 · 2010
Earlier work this paper cites.
Learning an Animatable Detailed 3D Face Model from In-The-Wild Images
Yao Feng, Haiwen Feng, Michael J. Black, and Timo Bolkart. 2021 · 2012
Earlier work this paper cites.
Animated speech: research progress and applications
D. W. Massaro, M. M. Cohen, M. Tabain, J. Beskow, and R. Clark. 2012 · 2012
Earlier work this paper cites.
JALI: an animator-centric viseme model for expressive lip synchronization
Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh. 2016 · 2016
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen. 2017 · 2017
Earlier work this paper cites.
Learning a model of facial shape and expression from 4D scans
Tianye Li, Timo Bolkart, Michael J. Black, Hao Li, and Javier Romero. 2017 · 2017
Earlier work this paper cites.
The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English
Steven R. Livingstone and Frank A. Russo. 2018 · 2018
Earlier work this paper cites.
Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis
Yuxuan Wang, Daisy Stanton, Yu Zhang, R. J. Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Fei Ren, Ye Jia, and Rif A. Saurous. 2018 · 2018
Earlier work this paper cites.
Learning Latent Representations for Style Control and Transfer in End-to-end Speech Synthesis
Ya-Jie Zhang, Shifeng Pan, Lei He, and Zhen-Hua Ling. 2019 · 2019
Earlier work this paper cites.
3D Morphable Face Models—Past, Present, and Future
Bernhard Egger, William A. P. Smith, Ayush Tewari, Stefanie Wuhrer, Michael Zollhoefer, Thabo Beeler, Florian Bernard, Timo Bolkart, Adam Kortylewski, Sami Romdhani, Christian Theobalt, Volker Blanz, and Thomas Vetter. 2020 · 2020
Cited alongside, same era.
MEAD: A Large-Scale Audio-Visual Dataset for Emotional Talking-Face Generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. 2020 · 2020
Cited alongside, same era.
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Transflower: probabilistic autoregressive dance generation with multimodal attention
Guillermo Valle-Pérez, Gustav Eje Henter, Jonas Beskow, André Holzapfel, Pierre-Yves Oudeyer, and Simon Alexanderson. 2021 · 2021
Cited alongside, same era.
CelebV-HQ: A Large-Scale Video Facial Attributes Dataset
Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. 2022 · 2022
Later among the works it cites.
Emotional speech-driven animation with content-emotion disentanglement. In SIGGRAPH Asia 2023 Conference Papers . 1–13
Radek Daněček, Kiran Chhatre, Shashank Tripathi, Yandong Wen, Michael Black, and Timo Bolkart. 2023 · 2023
Later among the works it cites.
Radek Daněček, Kiran Chhatre, Shashank Tripathi, Yandong Wen, Michael J. Black, and Timo Bolkart. 2023 · 2023
Later among the works it cites.
EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation
Ziqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu, Xiangyu Zhu, Jun He, Hongyan Liu, and Zhaoxin Fan. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Radek Daněček, Michael J. Black, and Timo Bolkart. 2022 · 2022
Cited alongside, same era.
FaceFormer: Speech-Driven 3D Facial Animation with Transformers
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura. 2022 · 2022
Cited alongside, same era.
ZeroEGGS: Zero-shot Example-based Gesture Generation from Speech
Saeed Ghorbani, Ylva Ferstl, Daniel Holden, Nikolaus F. Troje, and Marc-André Carbonneau. 2022 · 2022
Cited alongside, same era.
Classifier-Free Diffusion Guidance
Jonathan Ho and Tim Salimans. 2022 · 2022
Cited alongside, same era.
VOCAL: Vowel and Consonant Layering for Expressive Animator-Centric Singing Animation. In SIGGRAPH Asia 2022 Conference Papers (SA ’22) . Association for Computing Machinery, New York, NY, USA, 1–9
Yifang Pan, Chris Landreth, Eugene Fiume, and Karan Singh. 2022 · 2022
Cited alongside, same era.
MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement
Alexander Richard, Michael Zollhoefer, Yandong Wen, Fernando de la Torre, and Yaser Sheikh. 2022 · 2022
Cited alongside, same era.
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Cited alongside, same era.
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H. Bermano. 2022 · 2022
Cited alongside, same era.
CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Vancouver, BC, Canada, 12780–12790
Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong. 2023 · 2023
Later among the works it cites.
CelebV-Text: A Large-Scale Facial Text-Video Dataset
Jianhui Yu, Hao Zhu, Liming Jiang, Chen Change Loy, Weidong Cai, and Wayne Wu. 2023 · 2023
Later among the works it cites.
Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang, Xi Shen, Yu Guo, Ying Shan, and Fei Wang. 2023 · 2023
Later among the works it cites.
SEREP: Semantic Facial Expression Representation for Robust In-the-Wild Capture and Retargeting
Arthur Josi, Luiz Gustavo Hafemann, Abdallah Dib, Emeline Got, Rafael M. O. Cruz, and Marc-Andre Carbonneau. 2024 · 2024
Later among the works it cites.
Optimizing Diffusion Noise Can Serve As Universal Motion Priors
Korrawe Karunratanakul, Konpat Preechakul, Emre Aksan, Thabo Beeler, Supasorn Suwajanakorn, and Siyu Tang. 2024 · 2024
Later among the works it cites.
TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles
Yifeng Ma, Suzhen Wang, Yu Ding, Bowen Ma, Tangjie Lv, Changjie Fan, Zhipeng Hu, Zhidong Deng, and Xin Yu. 2024 · 2024
Later among the works it cites.
Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models
Zhiyao Sun, Tian Lv, Sheng Ye, Matthieu Lin, Jenny Sheng, Yu-Hui Wen, Minjing Yu, and Yong-jin Liu. 2024 · 2024
Later among the works it cites.
VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
Sicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang, Chong Li, Zhenyu Zang, Yizhong Zhang, Xin Tong, and Baining Guo. 2024 · 2024
Later among the works it cites.
SMGDiff: Soccer Motion Generation using diffusion probabilistic models
Hongdi Yang, Chengyang Li, Zhenxuan Wu, Gaozheng Li, Jingya Wang, Jingyi Yu, Zhuo Su, and Lan Xu. 2024 · 2024
Later among the works it cites.
Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance
Qingcheng Zhao, Pengyu Long, Qixuan Zhang, Dafei Qin, Han Liang, Longwen Zhang, Yingliang Zhang, Jingyi Yu, and Lan Xu. 2024 · 2024
Later among the works it cites.