Fetching the paper…
Reading the bibliography…
Speech-driven 3D facial animation synthesis has been a challenging task both in industry and research.
Expressive Speech-Driven Facial Animation
Yong Cao, Wen C. Tien, Petros Faloutsos, and Frédéric Pighin. 2005 · 2005
Earlier work this paper cites.
A 3-d audio-visual corpus of affective communication
Gabriele Fanelli, Juergen Gall, Harald Romsdorfer, Thibaut Weise, and Luc Van Gool. 2010 · 2010
Earlier work this paper cites.
Dynamic Units of Visual Speech. In Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation (Lausanne, Switzerland) (SCA ’12) . Eurographics Association, Goslar, DEU, 275–284
Sarah L. Taylor, Moshe Mahler, Barry-John Theobald, and Iain Matthews. 2012 · 2012
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al · 2014
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 5206–5210
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning . PMLR, 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen. 2017 · 2017
Earlier work this paper cites.
Learning a model of facial shape and expression from 4D scans
Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero. 2017 · 2017
Earlier work this paper cites.
A deep learning approach for generalized speech animation
Sarah Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews. 2017 · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc
Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu. 2017 · 2017
Earlier work this paper cites.
Visemenet: Audio-driven animator-centric speech animation
Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh. 2018 · 2018
Earlier work this paper cites.
Audio-driven emotional speech animation for interactive virtual characters
Constantinos Charalambous, Zerrin Yumak, and A.F. van der Stappen. 2019 · 2019
Earlier work this paper cites.
Capture, learning, and synthesis of 3D speaking styles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10101–10111
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Learning to Regress 3D Face Shape and Expression from an Image without 3D Supervision. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) . 7763–7772
Soubhik Sanyal, Timo Bolkart, Haiwen Feng, and Michael Black. 2019 · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Earlier work this paper cites.
3D Morphable Face Models—Past, Present, and Future
Bernhard Egger, William A. P. Smith, Ayush Tewari, Stefanie Wuhrer, Michael Zollhoefer, Thabo Beeler, Florian Bernard, Timo Bolkart, Adam Kortylewski, Sami Romdhani, Christian Theobalt, Volker Blanz, and Thomas Vetter. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020b · 2020
Earlier work this paper cites.
Let’s face it: Probabilistic multi-modal interlocutor-aware generation of facial gestures in dyadic settings. In International Conference on Intelligent Virtual Agents (IVA ’20) . ACM
Patrik Jonell, Taras Kucherenko, Gustav Eje Henter, and Jonas Beskow. 2020 · 2020
Cited alongside, same era.
Learning an Animatable Detailed 3D Face Model from In-The-Wild Images
Yao Feng, Haiwen Feng, Michael J. Black, and Timo Bolkart. 2021 · 2021
Cited alongside, same era.
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Live Speech Portraits: Real-Time Photorealistic Talking-Head Animation
Yuanxun Lu, Jinxiang Chai, and Xun Cao. 2021 · 2021
Cited alongside, same era.
Survey on 3D face reconstruction from uncalibrated images
Araceli Morales, Gemma Piella, and Federico M. Sukno. 2021 · 2021
Cited alongside, same era.
Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
Simon Alexanderson, Rajmund Nagy, Jonas Beskow, and Gustav Eje Henter. 2023 · 2023
Closest in time.
Speech Driven Video Editing via an Audio-Conditioned Diffusion Model
Dan Bigioi, Shubhajit Basak, Hugh Jordan, Rachel McDonnell, and Peter Corcoran. 2023 · 2023
Closest in time.
Diffusion models in vision: A survey
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah. 2023 · 2023
Closest in time.
Emotional Speech-Driven Animation with Content-Emotion Disentanglement
Radek Daněček, Kiran Chhatre, Shashank Tripathi, Yandong Wen, Michael J. Black, and Timo Bolkart. 2023 · 2023
Closest in time.
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
Chenpng Du, Qi Chen, Tianyu He, Xu Tan, Xie Chen, Kai Yu, Sheng Zhao, and Jiang Bian. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meshtalk: 3d face animation from speech using cross-modality disentanglement. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 1173–1182
Alexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre, and Yaser Sheikh. 2021 · 2021
Cited alongside, same era.
Score-Based Generative Modeling through Stochastic Differential Equations. In International Conference on Learning Representations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021 · 2021
Cited alongside, same era.
Voice2Face: Audio-driven Facial and Tongue Rig Animations with cVAEs. In EUROGRAPHICS SYMPOSIUM ON COMPUTER ANIMATION (SCA 2022
Mónica Villanueva Aylagas, Héctor Anadon Leon, Mattias Teye, and Konrad Tollmar. 2022 · 2022
Cited alongside, same era.
EMOCA: Emotion Driven Monocular Face Capture and Animation. In Conference on Computer Vision and Pattern Recognition (CVPR) . 20311–20322
Radek Danecek, Michael J. Black, and Timo Bolkart. 2022 · 2022
Cited alongside, same era.
Joint Audio-Text Model for Expressive Speech-Driven 3D Facial Animation
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura. 2022b · 2022
Cited alongside, same era.
EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion Model. In ACM SIGGRAPH 2022 Conference Proceedings (Vancouver, BC, Canada) (SIGGRAPH ’22) . Association for Computing Machinery, New York, NY, USA, Article 61, 10 pages
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Wayne Wu, Feng Xu, and Xun Cao. 2022 · 2022
Cited alongside, same era.
BEAT: A Large-Scale Semantic and Emotional Multi-Modal Dataset for Conversational Gestures Synthesis. In European conference on computer vision
Haiyang Liu, Zihao Zhu, Naoya Iwamoto, Yichen Peng, Zhengqing Li, You Zhou, Elif Bozkurt, and Bo Zheng. 2022 · 2022
Cited alongside, same era.
Closest in time.
Dynamixyz
Dynamixyz 2023 · 2023
Closest in time.
MetaHuman Animator
Epic Games 2023 · 2023
Closest in time.
Faceware
Faceware 2023 · 2023
Closest in time.
FaceXHuBERT: Text-less Speech-driven E(X)pressive 3D Facial Animation Synthesis Using Self-Supervised Speech Representation Learning. In INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI ’23) (Paris, France). ACM, New York, NY, USA, 10 pages
Kazi Injamamul Haque and Zerrin Yumak. 2023 · 2023
Closest in time.
JALI Research
JALI 2023 · 2023
Closest in time.
Prolific
Prolific 2023 · 2023
Closest in time.
Qualtrics
Qualtrics 2023 · 2023
Closest in time.
Ray Character Maya Scene by CGTarian
Ray CGTARIAN 2023 · 2023
Closest in time.
Edge: Editable dance generation from music. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 448–458
Jonathan Tseng, Rodrigo Castellon, and Karen Liu. 2023 · 2023
Closest in time.
CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior
Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong. 2023 · 2023
Closest in time.
DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models
Sicheng Yang, Zhiyong Wu, Minglei Li, Zhensong Zhang, Lei Hao, Weihong Bao, Ming Cheng, and Long Xiao. 2023 · 2023
Closest in time.
Generating Holistic 3D Human Motion from Speech. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 469–480
Hongwei Yi, Hualin Liang, Yifei Liu, Qiong Cao, Yandong Wen, Timo Bolkart, Dacheng Tao, and Michael J Black. 2023 · 2023
Closest in time.
SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8652–8661
Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang, Xi Shen, Yu Guo, Ying Shan, and Fei Wang. 2023 · 2023
Closest in time.