Fetching the paper…
Reading the bibliography…
We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal.
Marching cubes: A high resolution 3d surface construction algorithm
William E. Lorensen and Harvey E. Cline · 1987
Earlier work this paper cites.
Out of time: automated lip sync in the wild
J. S. Chung and A. Zisserman · 2016
Earlier work this paper cites.
Structure-from-motion revisited
Johannes Lutz Schönberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
You said that?, 2017
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman · 2017
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017
Stefan Elfwing, Eiji Uchibe, and Kenji Doya · 2017
Earlier work this paper cites.
The lj speech dataset
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Earlier work this paper cites.
Learning a model of facial shape and expression from 4D scans
Tianye Li, Timo Bolkart, Michael. J. Black, Hao Li, and Javier Romero · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer, 2017
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron Courville · 2017
Earlier work this paper cites.
Synthesizing obama: Learning lip sync from audio
Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman · 2017
Earlier work this paper cites.
Lip movements generation at a glance, 2018
Lele Chen, Zhiheng Li, Ross K. Maddox, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
End-to-end speech-driven facial animation with temporal gans, 2018
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2018
Earlier work this paper cites.
X2face: A network for controlling face generation by using images, audio, and pose codes
O. Wiles, A.S. Koepke, and A. Zisserman · 2018
Earlier work this paper cites.
Language2pose: Natural language grounded pose forecasting, 2019
Chaitanya Ahuja and Louis-Philippe Morency · 2019
Earlier work this paper cites.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Earlier work this paper cites.
Capture, learning, and synthesis of 3D speaking styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael Black · 2019
Earlier work this paper cites.
Pytorch lightning
William Falcon et al · 2019
Earlier work this paper cites.
Squeeze-and-excitation networks, 2019
Jie Hu, Li Shen, Samuel Albanie, Gang Sun, and Enhua Wu · 2019
Earlier work this paper cites.
Towards automatic face-to-face translation
Prajwal K R, Rudrabha Mukhopadhyay, Jerin Philip, Abhishek Jha, Vinay Namboodiri, and C V Jawahar · 2019
Earlier work this paper cites.
timsainb/noisereduce: v1.0, 2019
Tim Sainburg · 2019
Earlier work this paper cites.
Realistic speech-driven facial animation with gans, 2019
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2019
Earlier work this paper cites.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations, 2020
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Talking-head generation with rhythmic head motion
Lele Chen, Guofeng Cui, Celong Liu, Zhong Li, Ziyi Kou, Yi Xu, and Chenliang Xu · 2020
Earlier work this paper cites.
Speech-driven facial animation using cascaded gans for learning of motion and texture
Dipanjan Das, Sandika Biswas, Sanjana Sinha, and Brojeshwar Bhowmick · 2020
Earlier work this paper cites.
Action2motion: Conditioned generation of 3d human motions
Chuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, and Li Cheng · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Mish: A self regularized non-monotonic activation function, 2020
Diganta Misra · 2020
Earlier work this paper cites.
Pymcubes
pmneila · 2020
Cited alongside, same era.
A lip sync expert is all you need for speech to lip generation in the wild
K R Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, and C.V. Jawahar · 2020
Cited alongside, same era.
Everybody’s talkin’: Let me talk as you want, 2020
Linsen Song, Wayne Wu, Chen Qian, Ran He, and Chen Change Loy · 2020
Cited alongside, same era.
Neural voice puppetry: Audio-driven facial reenactment
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner · 2020
Cited alongside, same era.
Makelttalk: Speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li · 2020
Cited alongside, same era.
Efficient geometry-aware 3D generative adversarial networks
Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein · 2021
Learning dynamic facial radiance fields for few-shot talking head synthesis
Shuai Shen, Wanhua Li, Zheng Zhu, Yueqi Duan, Jie Zhou, and Jiwen Lu · 2022
Later among the works it cites.
Memories are one-to-many mapping alleviators in talking face generation
Anni Tang, Tianyu He, Xu Tan, Jun Ling, Runnan Li, Sheng Zhao, Li Song, and Jiang Bian · 2022
Later among the works it cites.
Edge: Editable dance generation from music
Jonathan Tseng, Rodrigo Castellon, and C Karen Liu · 2022
Later among the works it cites.
Mcvd: Masked conditional video diffusion for prediction, generation, and interpolation
Vikram Voleti, Alexia Jolicoeur-Martineau, and Christopher Pal · 2022
Later among the works it cites.
Dfa-nerf: Personalized talking head generation via disentangled face attributes neural rendering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Diffusion models beat gans on image synthesis, 2021
Prafulla Dhariwal and Alex Nichol · 2021
Cited alongside, same era.
Headgan: One-shot neural head synthesis and editing
Michail Christos Doukas, Stefanos Zafeiriou, and Viktoriia Sharmanska · 2021
Cited alongside, same era.
Dynamic neural radiance fields for monocular 4d facial avatar reconstruction
Guy Gafni, Justus Thies, Michael Zollhöfer, and Matthias Nießner · 2021
Cited alongside, same era.
Ad-nerf: Audio driven neural radiance fields for talking head synthesis
Yudong Guo, Keyu Chen, Sen Liang, Yongjin Liu, Hujun Bao, and Juyong Zhang · 2021
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2021
Cited alongside, same era.
Audio-driven emotional video portraits
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu · 2021
Cited alongside, same era.
Shunyu Yao, RuiZhe Zhong, Yichao Yan, Guangtao Zhai, and Xiaokang Yang · 2022
Later among the works it cites.
Listen, denoise, action! audio-driven motion synthesis with diffusion models
Simon Alexanderson, Rajmund Nagy, Jonas Beskow, and Gustav Eje Henter · 2023
Closest in time.
SINC: Spatial composition of 3D human motions for simultaneous action generation
Nikos Athanasiou, Mathis Petrovich, Michael J. Black, and Gül Varol · 2023
Closest in time.
Make-an-animation: Large-scale text-conditional 3d human motion generation, 2023
Samaneh Azadi, Akbar Shah, Thomas Hayes, Devi Parikh, and Sonal Gupta · 2023
Closest in time.
A morphable model for the synthesis of 3d faces
Volker Blanz and Thomas Vetter · 2023
Closest in time.
Align your latents: High-resolution video synthesis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis · 2023
Closest in time.
Emotional speech-driven animation with content-emotion disentanglement
Radek Daněček, Kiran Chhatre, Shashank Tripathi, Yandong Wen, Michael Black, and Timo Bolkart · 2023
Closest in time.
Blenderproc2: A procedural pipeline for photorealistic rendering
Maximilian Denninger, Dominik Winkelbauer, Martin Sundermeyer, Wout Boerdijk, Markus Knauer, Klaus H. Strobl, Matthias Humt, and Rudolph Triebel · 2023
Closest in time.
Sketching the future (stf): Applying conditional control techniques to text-to-video models, 2023
Rohan Dhesikan and Vignesh Rajmohan · 2023
Closest in time.
Hierarchical masked 3d diffusion model for video outpainting, 2023
Fanda Fan, Chaoxu Guo, Litong Gong, Biao Wang, Tiezheng Ge, Yuning Jiang, Chunjie Luo, and Jianfeng Zhan · 2023
Closest in time.
Learning neural parametric head models
Simon Giebenhain, Tobias Kirschstein, Markos Georgopoulos, Martin Rünz, Lourdes Agapito, and Matthias Nießner · 2023
Closest in time.
Nersemble: Multi-view radiance field reconstruction of human heads
Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, and Matthias Nießner · 2023
Closest in time.
Emotalk: Speech-driven emotional disentanglement for 3d face animation
Ziqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu, Xiangyu Zhu, Jun He, Hongyan Liu, and Zhaoxin Fan · 2023
Closest in time.
Diffusion motion: Generate text-guided 3d human motion by diffusion model, 2023
Zhiyuan Ren, Zhihong Pan, Xin Zhou, and Le Kang · 2023
Closest in time.
Difftalk: Crafting diffusion models for generalized audio-driven portraits animation
Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu · 2023
Closest in time.
Facediffuser: Speech-driven 3d facial animation synthesis using diffusion
Stefan Stan, Kazi Injamamul Haque, and Zerrin Yumak · 2023
Closest in time.
Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation
Michał Stypułkowski, Konstantinos Vougioukas, Sen He, Maciej Zikeba, Stavros Petridis, and Maja Pantic · 2023
Closest in time.
Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models, 2023
Zhiyao Sun, Tian Lv, Sheng Ye, Matthieu Gaetan Lin, Jenny Sheng, Yu-Hui Wen, Minjing Yu, and Yong jin Liu · 2023
Closest in time.
Human motion diffusion model
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano · 2023
Closest in time.
Imitator: Personalized speech-driven 3d facial animation
Balamurugan Thambiraja, Ikhsanul Habibie, Sadegh Aliakbarian, Darren Cosker, Christian Theobalt, and Justus Thies · 2023
Closest in time.
Attention is all you need, 2023
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2023
Closest in time.
Exploring video quality assessment on user generated contents from aesthetic and technical perspectives
Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou Hou, Annan Wang, Wenxiu Sun Sun, Qiong Yan, and Weisi Lin · 2023
Closest in time.
Codetalker: Speech-driven 3d facial animation with discrete motion prior
Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong · 2023
Closest in time.
Taming diffusion models for audio-driven co-speech gesture generation
Lingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian, Ziwei Liu, and Lequan Yu · 2023
Closest in time.
MonoNPHM: Dynamic head reconstruction from monocular videos
Simon Giebenhain, Tobias Kirschstein, Markos Georgopoulos, Martin Rünz, Lourdes Agapito, and Matthias Nießner · 2024
Closest in time.