Fetching the paper…
Reading the bibliography…
The generation of stylistic 3D facial animations driven by speech presents a significant challenge as it requires learning a many-to-many mapping between speech, style, and the corresponding natural facial motion.
Animated speech: research progress and applications. In AVSP . ISCA, Aalborg, Denmark, 200
Michael M. Cohen, Rashid Clark, and Dominic W. Massaro. 2001 · 2001
Earlier work this paper cites.
Random Forests for Real Time 3D Face Analysis
Gabriele Fanelli, Matthias Dantone, Juergen Gall, Andrea Fossati, and Luc Van Gool. 2013 · 2013
Earlier work this paper cites.
Deep Speech: Scaling up end-to-end speech recognition
Awni Y. Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, and Andrew Y. Ng. 2014 · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization. In ICLR (Poster) . ICLR, San Diego, CA., 1–15
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Deep Unsupervised Learning using Nonequilibrium Thermodynamics. In ICML (JMLR Workshop and Conference Proceedings, Vol. 37) . JMLR.org, Lille, France, 2256–2265
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
JALI: an animator-centric viseme model for expressive lip synchronization
Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh. 2016 · 2016
Earlier work this paper cites.
Learning a model of facial shape and expression from 4D scans
Tianye Li, Timo Bolkart, Michael J. Black, Hao Li, and Javier Romero. 2017 · 2017
Earlier work this paper cites.
Capture, Learning, and Synthesis of 3D Speaking Styles. In CVPR . Computer Vision Foundation / IEEE, Long Beach, CA, USA, 10101–10111
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J. Black. 2019 · 2019
Earlier work this paper cites.
3D Guided Fine-Grained Face Manipulation. In CVPR . Computer Vision Foundation / IEEE, Long Beach, CA, USA, 9821–9830
Zhenglin Geng, Chen Cao, and Sergey Tulyakov. 2019 · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Earlier work this paper cites.
A Simple Framework for Contrastive Learning of Visual Representations. In ICML (Proceedings of Machine Learning Research, Vol. 119) . PMLR, Virtual Event, 1597–1607
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. 2021 · 2021
Earlier work this paper cites.
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Earlier work this paper cites.
MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement. In ICCV . IEEE, Montreal, QC, Canada, 1153–1162
Alexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre, and Yaser Sheikh. 2021 · 2021
Cited alongside, same era.
Flow-Guided One-Shot Talking Face Generation With a High-Resolution Audio-Visual Dataset. In CVPR . Computer Vision Foundation / IEEE, Virtual Event, 3661–3670
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan. 2021 · 2021
Cited alongside, same era.
FaceFormer: Speech-Driven 3D Facial Animation with Transformers. In CVPR . IEEE, New Orleans, LA, USA, 18749–18758
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura. 2022 · 2022
Cited alongside, same era.
Visual Speech-Aware Perceptual 3D Facial Expression Reconstruction from Videos
Panagiotis Paraskevas Filntisis, George Retsinas, Foivos Paraperas Papantoniou, Athanasios Katsamanis, Anastasios Roussos, and Petros Maragos. 2022 · 2022
Cited alongside, same era.
InstructPix2Pix: Learning to Follow Image Editing Instructions. In CVPR . IEEE, Vancouver, BC, Canada, 18392–18402
Tim Brooks, Aleksander Holynski, and Alexei A. Efros. 2023 · 2023
Closest in time.
Emotional Speech-Driven Animation with Content-Emotion Disentanglement. In SIGGRAPH Asia 2023 Conference Papers (, Sydney, NSW, Australia,) (SA ’23) . Association for Computing Machinery, New York, NY, USA, Article 41, 13 pages
Radek Daněček, Kiran Chhatre, Shashank Tripathi, Yandong Wen, Michael Black, and Timo Bolkart. 2023 · 2023
Closest in time.
FaceXHuBERT: Text-less Speech-driven E(X)pressive 3D Facial Animation Synthesis Using Self-Supervised Speech Representation Learning. In ICMI . ACM, Paris, France, 282–291
Kazi Injamamul Haque and Zerrin Yumak. 2023 · 2023
Closest in time.
EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation. In Proceedings of the IEEE/CVF international conference on computer vision . IEEE, Vancouver, 20687–20697
Ziqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu, Xiangyu Zhu, Hongyan Liu, Jun He, and Zhaoxin Fan. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
6d Rotation Representation For Unconstrained Head Pose Estimation. In ICIP . IEEE, Bordeaux, France, 2496–2500
Thorsten Hempel, Ahmed A. Abdelrahman, and Ayoub Al-Hamadi. 2022 · 2022
Cited alongside, same era.
Classifier-Free Diffusion Guidance
Jonathan Ho and Tim Salimans. 2022 · 2022
Cited alongside, same era.
Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic Memory. In CVPR . IEEE, New Orleans, Louisiana, USA, 11040–11049
Siyao Li, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. 2022 · 2022
Cited alongside, same era.
DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022a · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022b · 2022
Cited alongside, same era.
High-Resolution Image Synthesis with Latent Diffusion Models. In CVPR . IEEE, New Orleans, LA, USA, 10674–10685
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Cited alongside, same era.
Predicting Personalized Head Movement From Short Video and Speech Signal
Ran Yi, Zipeng Ye, Zhiyao Sun, Juyong Zhang, Guoxin Zhang, Pengfei Wan, Hujun Bao, and Yong-Jin Liu. 2023 · 2022
Cited alongside, same era.
GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians
Shenhan Qian, Tobias Kirschstein, Liam Schoneveld, Davide Davoli, Simon Giebenhain, and Matthias Nießner. 2023 · 2023
Closest in time.
Diffusion motion: Generate text-guided 3d human motion by diffusion model. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, Rhodes, Greece, 1–5
Zhiyuan Ren, Zhihong Pan, Xin Zhou, and Le Kang. 2023 · 2023
Closest in time.
FaceDiffuser: Speech-Driven 3D Facial Animation Synthesis Using Diffusion. In MIG . ACM, Rennes, France, 13:1–13:11
Stefan Stan, Kazi Injamamul Haque, and Zerrin Yumak. 2023 · 2023
Closest in time.
Continuously Controllable Facial Expression Editing in Talking Face Videos
Zhiyao Sun, Yu-Hui Wen, Tian Lv, Yanan Sun, Ziyang Zhang, Yaoyuan Wang, and Yong-Jin Liu. 2023 · 2023
Closest in time.
Human Motion Diffusion Model. In The Eleventh International Conference on Learning Representations . https://openreview.net/forum?id=SJ1kSyO2jwu, Kigali, Rwanda, 1–12
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. 2023 · 2023
Closest in time.
Imitator: Personalized speech-driven 3d facial animation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . IEEE, Vancouver, 20621–20631
Balamurugan Thambiraja, Ikhsanul Habibie, Sadegh Aliakbarian, Darren Cosker, Christian Theobalt, and Justus Thies. 2023 · 2023
Closest in time.
CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior. In CVPR . IEEE, Vancouver, 12780–12790
Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong. 2023 · 2023
Closest in time.
Diffusion models: A comprehensive survey of methods and applications
Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2023 · 2023
Closest in time.
3D Talking Face With Personalized Pose Dynamics
Chenxu Zhang, Saifeng Ni, Zhipeng Fan, Hongbo Li, Ming Zeng, Madhukar Budagavi, and Xiaohu Guo. 2023b · 2023
Closest in time.
Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation. In CVPR . IEEE, Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation, 10544–10553
Lingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian, Ziwei Liu, and Lequan Yu. 2023 · 2023
Closest in time.