Fetching the paper…
Reading the bibliography…
Recent studies in speech-driven 3D talking head generation have achieved convincing results in verbal articulations.
H. McGurk and J. MacDonald, “Hearing lips and seeing voices,” Nature , 1976
1976
Earlier work this paper cites.
C. Liu, “An analysis of the current and future state of 3d facial animation techniques and systems,” Simon Fraser University , 2009
2009
Earlier work this paper cites.
G. Fanelli, J. Gall, H. Romsdorfer, T. Weise, and L. Van Gool, “A 3-d audio-visual corpus of affective communication,” IEEE Transactions on Multimedia , 2010
2010
Earlier work this paper cites.
2013
Earlier work this paper cites.
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen, “Audio-driven facial animation by joint end-to-end learning of pose and emotion,” ACM Transactions on Graphics (SIGGRAPH) , 2017
2017
Earlier work this paper cites.
T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero, “Learning a model of facial shape and expression from 4d scans.” ACM Transactions on Graphics (SIGGRAPH) , 2017
2017
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” in Advances in Neural Information Processing Systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,” in Conference of the International Speech Communication Association (INTERSPEECH) , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in Conference of the International Speech Communication Association (INTERSPEECH) , 2018
2018
Earlier work this paper cites.
D. Cudeiro, T. Bolkart, C. Laidlaw, A. Ranjan, and M. J. Black, “Capture, learning, and synthesis of 3d speaking styles,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
I. Wohlgenannt, A. Simons, and S. Stieglitz, “Virtual reality,” Business & Information Systems Engineering , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Richard, M. Zollhöfer, Y. Wen, F. De la Torre, and Y. Sheikh, “Meshtalk: 3d face animation from speech using cross-modality disentanglement,” in IEEE International Conference on Computer Vision (ICCV) , 2021
2021
Cited alongside, same era.
Z. Peng, H. Wu, Z. Song, H. Xu, X. Zhu, H. Liu, J. He, and Z. Fan, “Emotalk: Speech-driven emotional disentanglement for 3d face animation,” in IEEE International Conference on Computer Vision (ICCV) , 2023
2023
Later among the works it cites.
R. Daněček, K. Chhatre, S. Tripathi, Y. Wen, M. J. Black, and T. Bolkart, “Emotional speech-driven animation with content-emotion disentanglement,” in ACM Transactions on Graphics (SIGGRAPH Asia) , 2023
2023
Later among the works it cites.
P. P. Filntisis, G. Retsinas, F. Paraperas-Papantoniou, A. Katsamanis, A. Roussos, and P. Maragos, “Spectre: Visual speech-informed perceptual 3d facial expression reconstruction from videos,” in IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , 2023
2023
Later among the works it cites.
Z. Peng, Y. Luo, Y. Shi, H. Xu, X. Zhu, H. Liu, J. He, and Z. Fan, “Selftalk: A self-supervised commutative training diagram to comprehend 3d talking faces,” in ACM International Conference on Multimedia (MM) , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Huang, Z. Wu, S. Kang, D. Dai, J. Jia, T. Fu, D. Tuo, G. Lei, P. Liu, D. Su et al. , “Speaker independent and multilingual/mixlingual speech-driven talking head generation using phonetic posteriorgrams,” in 2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) , 2021
2021
Cited alongside, same era.
R. Tao, Z. Pan, R. K. Das, X. Qian, M. Z. Shou, and H. Li, “Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection,” in ACM International Conference on Multimedia (MM) , 2021
2021
Cited alongside, same era.
Y. Fan, Z. Lin, J. Saito, W. Wang, and T. Komura, “Faceformer: Speech-driven 3d facial animation with transformers,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2022
2022
Cited alongside, same era.
B. Thambiraja, I. Habibie, S. Aliakbarian, D. Cosker, C. Theobalt, and J. Thies, “Imitator: Personalized speech-driven 3d facial animation,” in IEEE International Conference on Computer Vision (ICCV) , 2022
2022
Cited alongside, same era.
E. Ng, H. Joo, L. Hu, H. Li, T. Darrell, A. Kanazawa, and S. Ginosar, “Learning to listen: Modeling non-deterministic dyadic facial motion,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2022
2022
Cited alongside, same era.
C. Sondermann and M. Merkt, “Like it or learn from it: Effects of talking heads in educational videos,” Computers & Education , 2023
2023
Cited alongside, same era.
J. Xing, M. Xia, Y. Zhang, X. Cun, J. Wang, and T.-T. Wong, “Codetalker: Speech-driven 3d facial animation with discrete motion prior,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
M. Anwar, B. Shi, V. Goswami, W.-N. Hsu, J. Pino, and C. Wang, “Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation,” in Conference of the International Speech Communication Association (INTERSPEECH) , 2023
2023
Later among the works it cites.
J. Yu, H. Zhu, L. Jiang, C. C. Loy, W. Cai, and W. Wu, “Celebv-text: A large-scale facial text-video dataset,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2023
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International Conference on Machine Learning (ICML) , 2023
2023
Later among the works it cites.
P. Ma, A. Haliassos, A. Fernandez-Lopez, H. Chen, S. Petridis, and M. Pantic, “Auto-avsr: Audio-visual speech recognition with automatic labels,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2023
2023
Later among the works it cites.
K. Sung-Bin, L. Hyun, D. H. Hong, S. Nam, J. Ju, and T.-H. Oh, “Laughtalk: Expressive 3d talking head generation with laughter,” in IEEE Winter Conference on Applications of Computer Vision (WACV) , 2024
2024
Closest in time.
J. H. Yeo, M. Kim, S. Watanabe, and Y. M. Ro, “Visual speech recognition for low-resource languages with automatic labels from whisper model,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2024
2024
Closest in time.