Fetching the paper…
Reading the bibliography…
To be widely adopted, 3D facial avatars must be animated easily, realistically, and directly from speech signals.
Animated speech: research progress and applications
Michael M. Cohen, Rashid Clark, and Dominic W. Massaro · 2001
Earlier work this paper cites.
Expressive speech-driven facial animation
Yong Cao, Wen C. Tien, Petros Faloutsos, and Frédéric Pighin · 2005
Earlier work this paper cites.
A 3-d audio-visual corpus of affective communication
G. Fanelli, J. Gall, H. Romsdorfer, T. Weise, and L. Van Gool · 2010
Earlier work this paper cites.
Dynamic units of visual speech
Sarah L. Taylor, Moshe Mahler, Barry-John Theobald, and Iain A. Matthews · 2012
Earlier work this paper cites.
A practical and configurable lip sync method for games
Yuyu Xu, Andrew W. Feng, Stacy Marsella, and Ari Shapiro · 2013
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David Cooper, Michael Keutmann, Ruben Gur, Ani Nenkova, and Ragini Verma · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
A review of eye gaze in virtual agents, social robotics and hci: Behaviour generation, user interaction and perception
K. Ruhland, C. E. Peters, S. Andrist, J. B. Badler, N. I. Badler, M. Gleicher, B. Mutlu, and R. McDonnell · 2015
Earlier work this paper cites.
Jali: An animator-centric viseme model for expressive lip synchronization
Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
MOSI: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos
Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency · 2016
Earlier work this paper cites.
Bringing portraits to life
Hadar Averbuch-Elor, Daniel Cohen-Or, Johannes Kopf, and Michael F. Cohen · 2017
Earlier work this paper cites.
How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230, 000 3d facial landmarks)
Adrian Bulat and Georgios Tzimiropoulos · 2017
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Earlier work this paper cites.
Learning a model of facial shape and expression from 4D scans
Tianye Li, Timo Bolkart, Michael. J. Black, Hao Li, and Javier Romero · 2017
Earlier work this paper cites.
Affectnet: A database for facial expression, valence, and arousal computing in the wild
Ali Mollahosseini, Behzad Hasani, and Mohammad H Mahoor · 2017
Earlier work this paper cites.
Voxceleb: A large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman · 2017
Earlier work this paper cites.
Speech-driven 3d facial animation with implicit emotional awareness: A deep learning approach
Hai Xuan Pham, Samuel Cheung, and Vladimir Pavlovic · 2017
Earlier work this paper cites.
End-to-end learning for 3d facial animation from raw waveforms of speech
Hai Xuan Pham, Yuting Wang, and Vladimir Pavlovic · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman · 2017
Earlier work this paper cites.
A deep learning approach for generalized speech animation
Sarah L. Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica K. Hodgins, and Iain A. Matthews · 2017
Earlier work this paper cites.
Mofa: Model-based deep convolutional face autoencoder for unsupervised monocular reconstruction
Ayush Tewari, Michael Zollhöfer, Hyeongwoo Kim, Pablo Garrido, Florian Bernard, Patrick Pérez, and Christian Theobalt · 2017
Earlier work this paper cites.
LRS3-TED: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Earlier work this paper cites.
Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Vggface2: A dataset for recognising faces across pose and age
Qiong Cao, Li Shen, Weidi Xie, Omkar M. Parkhi, and Andrew Zisserman · 2018
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Cited alongside, same era.
Exprgan: Facial expression editing with controllable expression intensity
Hui Ding, Kumar Sricharan, and Rama Chellappa · 2018
Cited alongside, same era.
Unsupervised training for 3d morphable model regression
Kyle Genova, Forrester Cole, Aaron Maschinot, Aaron Sarna, Daniel Vlasic, and William T. Freeman · 2018
Cited alongside, same era.
The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english
Steven R Livingstone and Frank A Russo · 2018
Cited alongside, same era.
Faceforensics: A large-scale video dataset for forgery detection in human faces
Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner · 2018
Cited alongside, same era.
Self-supervised monocular 3d face reconstruction by occlusion-aware multi-view geometry consistency
Jiaxiang Shang, Tianwei Shen, Shiwei Li, Lei Zhou, Mingmin Zhen, Tian Fang, and Long Quan · 2020
Later among the works it cites.
Neural voice puppetry: Audio-driven facial reenactment
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner · 2020
Later among the works it cites.
Icface: Interpretable and controllable face reenactment using gans
Soumya Tripathy, Juho Kannala, and Esa Rahtu · 2020
Later among the works it cites.
MEAD: A large-scale audio-visual dataset for emotional talking-face generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy · 2020
Later among the works it cites.
Facescape: A large-scale high quality 3d face dataset and detailed riggable 3d face prediction
Haotian Yang, Hao Zhu, Yanru Wang, Mingkai Huang, Qiu Shen, Ruigang Yang, and Xun Cao · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-supervised multi-level face model learning for monocular reconstruction at over 250 hz
Ayush Tewari, Michael Zollhöfer, Pablo Garrido, Florian Bernard, Hyeongwoo Kim, Patrick Pérez, and Christian Theobalt · 2018
Cited alongside, same era.
Visemenet: Audio-driven animator-centric speech animation
Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh · 2018
Cited alongside, same era.
State of the art on monocular 3d face reconstruction, tracking, and applications
Michael Zollhöfer, Justus Thies, Pablo Garrido, Derek Bradley, Thabo Beeler, Patrick Pérez, Marc Stamminger, Matthias Nießner, and Christian Theobalt · 2018
Cited alongside, same era.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K. Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Cited alongside, same era.
Capture, learning, and synthesis of 3d speaking styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J. Black · 2019
Cited alongside, same era.
Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set
Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
GANmut: Learning interpretable conditional space for gamut of emotions
Stefano d’Apolito, Danda Pani Paudel, Zhiwu Huang, Andres Romero, and Luc Van Gool · 2021
Later among the works it cites.
Learning an animatable detailed 3D face model from in-the-wild images
Yao Feng, Haiwen Feng, Michael J. Black, and Timo Bolkart · 2021
Later among the works it cites.
Audio-driven emotional video portraits
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu · 2021
Later among the works it cites.
Meshtalk: 3d face animation from speech using cross-modality disentanglement
Alexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre, and Yaser Sheikh · 2021
Later among the works it cites.
FACEGAN: facial attribute controllable reenactment GAN
Soumya Tripathy, Juho Kannala, and Esa Rahtu · 2021
Later among the works it cites.
Deep audio-visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew W. Senior, Oriol Vinyals, and Andrew Zisserman · 2022
Later among the works it cites.
Facial animation with disentangled identity and motion using transformers
Prashanth Chandran, Gaspard Zoss, Markus H. Gross, Paulo F. U. Gotardo, and Derek Bradley · 2022
Later among the works it cites.
Videoretalking: Audio-based lip synchronization for talking head video editing in the wild
Kun Cheng, Xiaodong Cun, Yong Zhang, Menghan Xia, Fei Yin, Mingrui Zhu, Xuan Wang, Jue Wang, and Nannan Wang · 2022
Later among the works it cites.
EMOCA: emotion driven monocular face capture and animation
Radek Danecek, Michael J. Black, and Timo Bolkart · 2022
Later among the works it cites.
Faceformer: Speech-driven 3d facial animation with transformers
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura · 2022
Later among the works it cites.
Visual speech-aware perceptual 3d facial expression reconstruction from videos, 2022
Panagiotis P. Filntisis, George Retsinas, Foivos Paraperas-Papantoniou, Athanasios Katsamanis, Anastasios Roussos, and Petros Maragos · 2022
Later among the works it cites.
Learning to listen: Modeling non-deterministic dyadic facial motion
Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li, Trevor Darrell, Angjoo Kanazawa, and Shiry Ginosar · 2022
Later among the works it cites.
Neural emotion director: Speech-preserving semantic control of facial expressions in ”in-the-wild” videos
Foivos Paraperas Papantoniou, Panagiotis Paraskevas Filntisis, Petros Maragos, and Anastasios Roussos · 2022
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A. Smith, and Mike Lewis · 2022
Later among the works it cites.
Multiface: A dataset for neural face rendering
Cheng-hsin Wuu, Ningyuan Zheng, Scott Ardisson, Rohan Bali, Danielle Belko, Eric Brockmeyer, Lucas Evans, Timothy Godisart, Hyowon Ha, Alexander Hypes, Taylor Koska, Steven Krenn, Stephen Lombardi, Xiaomin Luo, Kevyn McPhail, Laura Millerschoen, Michal Perdoch, Mark Pitts, Alexander Richard, Jason M. Saragih, Junko Saragih, Takaaki Shiratori, Tomas Simon, Matt Stewart, Autumn Trimble, Xinshuo Weng, David Whitewolf, Chenglei Wu, Shoou-I Yu, and Yaser Sheikh · 2022
Later among the works it cites.
Celebv-hq: A large-scale video facial attributes dataset
Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy · 2022
Later among the works it cites.
Towards metrical reconstruction of human faces
Wojciech Zielonka, Timo Bolkart, and Justus Thies · 2022
Later among the works it cites.
Emotalk: Speech-driven emotional disentanglement for 3d face animation
Ziqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu, Xiangyu Zhu, Hongyan Liu, Jun He, and Zhaoxin Fan · 2023
Closest in time.
Imitator: Personalized speech-driven 3d facial animation, 2023
Balamurugan Thambiraja, Ikhsanul Habibie, Sadegh Aliakbarian, Darren Cosker, Christian Theobalt, and Justus Thies · 2023
Closest in time.
Codetalker: Speech-driven 3d facial animation with discrete motion prior
Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong · 2023
Closest in time.