Fetching the paper…
Reading the bibliography…
Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Emotion recognition by speech signals
Oh-Wook Kwon, Kwokleung Chan, Jiucang Hao, and Te-Won Lee · 2003
Earlier work this paper cites.
Speech emotion recognition using hidden markov models
Tin Lay Nwe, Say Wei Foo, and Liyanage C De Silva · 2003
Earlier work this paper cites.
Hidden markov model-based speech emotion recognition
Björn Schuller, Gerhard Rigoll, and Manfred Lang · 2003
Earlier work this paper cites.
Expressive speech-driven facial animation
Yong Cao, Wen C Tien, Petros Faloutsos, and Frédéric Pighin · 2005
Earlier work this paper cites.
Dynamic time warping
Meinard Müller · 2007
Earlier work this paper cites.
Automatic recognition of emotions from speech: a review of the literature and recommendations for practical realisation
Thurid Vogt, Elisabeth André, and Johannes Wagner · 2008
Earlier work this paper cites.
An analysis of the current and future state of 3d facial animation techniques and systems
Chen Liu · 2009
Earlier work this paper cites.
Lindasalwa Muda, Mumtaj Begam, and Irraivan Elamvazuthi · 2010
Earlier work this paper cites.
Expression transfer: A system to build 3d blend shapes for facial animation
Chandan Pawaskar, Wan-Chun Ma, Kieran Carnegie, John P Lewis, and Taehyun Rhee · 2013
Earlier work this paper cites.
Computer facial animation: A review
Heng Yu Ping, Lili Nurliyana Abdullah, Puteri Suhaiza Sulaiman, and Alfian Abdul Halin · 2013
Earlier work this paper cites.
A narrative literature review of games, animations and simulations to teach research methods and statistics
Elizabeth A Boyle, Ewan W MacArthur, Thomas M Connolly, Thomas Hainey, Madalina Manea, Anne Kärki, and Peter Van Rosmalen · 2014
Earlier work this paper cites.
Speech emotion recognition using cnn
Zhengwei Huang, Ming Dong, Qirong Mao, and Yongzhao Zhan · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Practice and theory of blendshape facial models
John P Lewis, Ken Anjyo, Taehyun Rhee, Mengjie Zhang, Frederic H Pighin, and Zhigang Deng · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Temporal convolutional networks: A unified approach to action segmentation
Colin Lea, Rene Vidal, Austin Reiter, and Gregory D Hager · 2016
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Earlier work this paper cites.
Learning a model of facial shape and expression from 4d scans
Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero · 2017
Earlier work this paper cites.
Disentangled representation learning gan for pose-invariant face recognition
Luan Tran, Xi Yin, and Xiaoming Liu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english
Steven R Livingstone and Frank A Russo · 2018
Cited alongside, same era.
A survey on deep learning: Algorithms, techniques, and applications
Samira Pouyanfar, Saad Sadiq, Yilin Yan, Haiman Tian, Yudong Tao, Maria Presa Reyes, Mei-Ling Shyu, Shu-Ching Chen, and Sundaraja S Iyengar · 2018
Cited alongside, same era.
Speech emotion recognition: Two decades in a nutshell, benchmarks, and ongoing trends
Björn W Schuller · 2018
Cited alongside, same era.
Speech emotion recognition using spectrogram & phoneme embedding
Promod Yenigalla, Abhay Kumar, Suraj Tripathi, Chirag Singh, Sibsambhu Kar, and Jithendra Vepa · 2018
Cited alongside, same era.
Geometry-guided dense perspective network for speech-driven facial animation
Jingying Liu, Binyuan Hui, Kun Li, Yunke Liu, Yu-Kun Lai, Yuxiang Zhang, Yebin Liu, and Jingyu Yang · 2021
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis · 2021
Later among the works it cites.
Meshtalk: 3d face animation from speech using cross-modality disentanglement
Alexander Richard, Michael Zollhöfer, Yandong Wen, Fernando de la Torre, and Yaser Sheikh · 2021
Later among the works it cites.
Facial: Synthesizing dynamic talking face with implicit attribute learning
Chenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng, Saifeng Ni, Madhukar Budagavi, and Xiaohu Guo · 2021
Later among the works it cites.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Cited alongside, same era.
Capture, learning, and synthesis of 3d speaking styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black · 2019
Cited alongside, same era.
Emotion recognition using deep learning approach from audio–visual emotional big data
M Shamim Hossain and Ghulam Muhammad · 2019
Cited alongside, same era.
Audio2face: Generating speech/face animation from single audio with attention-based bidirectional lstm networks
Guanzhong Tian, Yi Yuan, and Yong Liu · 2019
Cited alongside, same era.
Few-shot adversarial learning of realistic neural talking head models
Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
Modality dropout for improved performance-driven talking faces
Ahmed Hussen Abdelaziz, Barry-John Theobald, Paul Dixon, Reinhard Knothe, Nicholas Apostoloff, and Sachin Kajareker · 2020
Cited alongside, same era.
Vaw-gan for disentanglement and recomposition of emotional elements in speech
Kun Zhou, Berrak Sisman, and Haizhou Li · 2021
Later among the works it cites.
Disentangling audio content and emotion with adaptive instance normalization for expressive facial animation synthesis
Che-Jui Chang, Long Zhao, Sen Zhang, and Mubbasir Kapadia · 2022
Later among the works it cites.
Transformer-s2a: Robust and efficient speech-to-animation
Liyang Chen, Zhiyong Wu, Jun Ling, Runnan Li, Xu Tan, and Sheng Zhao · 2022
Later among the works it cites.
Faceformer: Speech-driven 3d facial animation with transformers
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura · 2022
Later among the works it cites.
Reconstruction-aware prior distillation for semi-supervised point cloud completion
Zhaoxin Fan, Yulin He, Zhicheng Wang, Kejian Wu, Hongyan Liu, and Jun He · 2022
Later among the works it cites.
Object level depth reconstruction for category level 6d object pose estimation from monocular rgb image
Zhaoxin Fan, Zhenbo Song, Jian Xu, Zhicheng Wang, Kejian Wu, Hongyan Liu, and Jun He · 2022
Later among the works it cites.
Deep learning on monocular object pose detection and tracking: A comprehensive overview
Zhaoxin Fan, Yazhi Zhu, Yulin He, Qi Sun, Hongyan Liu, and Jun He · 2022
Later among the works it cites.
Depth-aware generative adversarial network for talking head video generation
Fa-Ting Hong, Longhao Zhang, Li Shen, and Dan Xu · 2022
Later among the works it cites.
Neural emotion director: Speech-preserving semantic control of facial expressions in” in-the-wild” videos
Foivos Paraperas Papantoniou, Panagiotis P Filntisis, Petros Maragos, and Anastasios Roussos · 2022
Later among the works it cites.
Deep learning for visual speech analysis: A survey
Changchong Sheng, Gangyao Kuang, Liang Bai, Chenping Hou, Yulan Guo, Xin Xu, Matti Pietikäinen, and Li Liu · 2022
Later among the works it cites.
Continuously controllable facial expression editing in talking face videos
Zhiyao Sun, Yu-Hui Wen, Tian Lv, Yanan Sun, Ziyang Zhang, Yaoyuan Wang, and Yong-Jin Liu · 2022
Later among the works it cites.
Perceiving and modeling density for image dehazing
Tian Ye, Yunchen Zhang, Mingchao Jiang, Liang Chen, Yun Liu, Sixiang Chen, and Erkang Chen · 2022
Later among the works it cites.
Guangyan Zhang, Ying Qin, Wenjie Zhang, Jialun Wu, Mei Li, Yutao Gai, Feijun Jiang, and Tan Lee · 2022
Later among the works it cites.
Msp-former: Multi-scale projection transformer for single image desnowing
Sixiang Chen, Tian Ye, Yun Liu, Taodong Liao, Jingxia Jiang, Erkang Chen, and Peng Chen · 2023
Closest in time.