Fetching the paper…
Reading the bibliography…
Speech-driven 3D facial animation has gained significant attention for its ability to create realistic and expressive facial animations in 3D space based on speech.
Expressive speech-driven facial animation
Cao, Y.; Tien, W. C.; Faloutsos, P.; and Pighin, F. 2005 · 2005
Earlier work this paper cites.
A 3D facial expression database for facial behavior research
Yin, L.; Wei, X.; Sun, Y.; Wang, J.; and Rosato, M. J. 2006 · 2006
Earlier work this paper cites.
Bosphorus database for 3D face analysis
Savran, A.; Alyüz, N.; Dibeklioğlu, H.; Çeliktutan, O.; Gökberk, B.; Sankur, B.; and Akarun, L. 2008 · 2008
Earlier work this paper cites.
A 3D face model for pose and illumination invariant face recognition
Paysan, P.; Knothe, R.; Amberg, B.; Romdhani, S.; and Vetter, T. 2009 · 2009
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Facewarehouse: A 3d facial expression database for visual computing
Cao, C.; Weng, Y.; Zhou, S.; Tong, Y.; and Zhou, K. 2013 · 2013
Earlier work this paper cites.
A high-resolution spontaneous 3d dynamic facial expression database
Zhang, X.; Yin, L.; Cohn, J. F.; Canavan, S.; Reale, M.; Horowitz, A.; and Liu, P. 2013 · 2013
Earlier work this paper cites.
A 3D dynamic database for unconstrained face recognition
Alashkar, T.; Amor, B. B.; Daoudi, M.; and Berretti, S. 2014 · 2014
Earlier work this paper cites.
Bp4d-spontaneous: a high-resolution spontaneous 3d dynamic facial expression database
Zhang, X.; Yin, L.; Cohn, J. F.; Canavan, S.; Reale, M.; Horowitz, A.; Liu, P.; and Girard, J. M. 2014 · 2014
Earlier work this paper cites.
SMPL: A skinned multi-person linear model
Loper, M.; Mahmood, N.; Romero, J.; Pons-Moll, G.; and Black, M. J. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; and Ganguli, S. 2015 · 2015
Earlier work this paper cites.
Lip reading in the wild
Chung, J. S.; and Zisserman, A. 2017 · 2016
Earlier work this paper cites.
Multimodal spontaneous emotion corpus for human behavior analysis
Zhang, Z.; Girard, J. M.; Wu, Y.; Zhang, X.; Liu, P.; Ciftci, U.; Canavan, S.; Reale, M.; Horowitz, A.; Yang, H.; et al. 2016 · 2016
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Karras, T.; Aila, T.; Laine, S.; Herva, A.; and Lehtinen, J. 2017 · 2017
Earlier work this paper cites.
Voxceleb: a large-scale speaker identification dataset
Nagrani, A.; Chung, J. S.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
A deep learning approach for generalized speech animation
Taylor, S.; Kim, T.; Yue, Y.; Mahler, M.; Krahe, J.; Rodriguez, A. G.; Hodgins, J.; and Matthews, I. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Deep audio-visual speech recognition
Afouras, T.; Chung, J. S.; Senior, A.; Vinyals, O.; and Zisserman, A. 2018 · 2018
Earlier work this paper cites.
LRS3-TED: a large-scale dataset for visual speech recognition
Afouras, T.; Chung, J. S.; and Zisserman, A. 2018 · 2018
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
Chung, J. S.; Nagrani, A.; and Zisserman, A. 2018 · 2018
Earlier work this paper cites.
End-to-end learning for 3d facial animation from speech
Pham, H. X.; Wang, Y.; and Pavlovic, V. 2018 · 2018
Cited alongside, same era.
Generating 3D faces using convolutional mesh autoencoders
Ranjan, A.; Bolkart, T.; Sanyal, S.; and Black, M. J. 2018 · 2018
Cited alongside, same era.
Visemenet: Audio-driven animator-centric speech animation
Zhou, Y.; Xu, Z.; Landreth, C.; Kalogerakis, E.; Maji, S.; and Singh, K. 2018 · 2018
Cited alongside, same era.
Capture, learning, and synthesis of 3D speaking styles
Cudeiro, D.; Bolkart, T.; Laidlaw, C.; Ranjan, A.; and Black, M. J. 2019 · 2019
Cited alongside, same era.
Learning to regress 3D face shape and expression from an image without 3D supervision
Sanyal, S.; Bolkart, T.; Feng, H.; and Black, M. J. 2019 · 2019
Cited alongside, same era.
3d visual passcode: Speech-driven 3d facial dynamics for behaviometrics
Zhang, J.; and Fisher, R. B. 2019 · 2019
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Zhou, H.; Sun, Y.; Wu, W.; Loy, C. C.; Wang, X.; and Liu, Z. 2021 · 2021
Later among the works it cites.
FacesCape: 3D facial dataset and benchmark for single-view 3D face reconstruction
Zhu, H.; Yang, H.; Guo, L.; Zhang, Y.; Wang, Y.; Huang, M.; Shen, Q.; Yang, R.; and Cao, X. 2021 · 2021
Later among the works it cites.
Classifier-free diffusion guidance
Ho, J.; and Salimans, T. 2022 · 2022
Later among the works it cites.
Depth-aware generative adversarial network for talking head video generation
Hong, F.-T.; Zhang, L.; Shen, L.; and Xu, D. 2022 · 2022
Later among the works it cites.
Expressive talking head generation with granular audio-visual control
Liang, B.; Pan, Y.; Guo, Z.; Zhou, H.; Hong, Z.; Han, X.; Han, J.; Liu, J.; Ding, E.; and Wang, J. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Talking face generation by adversarially disentangled audio-visual representation
Zhou, H.; Liu, Y.; Liu, Z.; Luo, P.; and Wang, X. 2019 · 2019
Cited alongside, same era.
GIF: Generative interpretable faces
Ghosh, P.; Gupta, P. S.; Uziel, R.; Ranjan, A.; Black, M. J.; and Bolkart, T. 2020 · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Cited alongside, same era.
A lip sync expert is all you need for speech to lip generation in the wild
Prajwal, K.; Mukhopadhyay, R.; Namboodiri, V. P.; and Jawahar, C. 2020 · 2020
Cited alongside, same era.
Mead: A large-scale audio-visual dataset for emotional talking-face generation
Wang, K.; Wu, Q.; Song, L.; Yang, Z.; Wu, W.; Qian, C.; He, R.; Qiao, Y.; and Loy, C. C. 2020 · 2020
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P.; and Nichol, A. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Pretrained Diffusion Models for Unified Human Motion Synthesis
Ma, J.; Bai, S.; and Zhou, C. 2022 · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Later among the works it cites.
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
Tang, A.; He, T.; Tan, X.; Ling, J.; Li, R.; Zhao, S.; Song, L.; and Bian, J. 2022 · 2022
Later among the works it cites.
Tevet, G.; Raab, S.; Gordon, B.; Shafir, Y.; Cohen-Or, D.; and Bermano, A. H. 2022 · 2022
Later among the works it cites.
Multiface: A Dataset for Neural Face Rendering
Wuu, C.-h.; Zheng, N.; Ardisson, S.; Bali, R.; Belko, D.; Brockmeyer, E.; Evans, L.; Godisart, T.; Ha, H.; Hypes, A.; Koska, T.; Krenn, S.; Lombardi, S.; Luo, X.; McPhail, K.; Millerschoen, L.; Perdoch, M.; Pitts, M.; Richard, A.; Saragih, J.; Saragih, J.; Shiratori, T.; Simon, T.; Stewart, M.; Trimble, A.; Weng, X.; Whitewolf, D.; Wu, C.; Yu, S.-I.; and Sheikh, Y. 2022 · 2022
Later among the works it cites.
Motiondiffuse: Text-driven human motion generation with diffusion model
Zhang, M.; Cai, Z.; Pan, L.; Hong, F.; Guo, X.; Yang, L.; and Liu, Z. 2022 · 2022
Later among the works it cites.
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
Du, C.; Chen, Q.; He, T.; Tan, X.; Chen, X.; Yu, K.; Zhao, S.; and Bian, J. 2023 · 2023
Closest in time.
StyleTalk: One-shot Talking Head Generation with Controllable Speaking Styles
Ma, Y.; Wang, S.; Hu, Z.; Fan, C.; Lv, T.; Ding, Y.; Deng, Z.; and Yu, X. 2023 · 2023
Closest in time.
DiffTalk: Crafting Diffusion Models for Generalized Talking Head Synthesis
Shen, S.; Zhao, W.; Meng, Z.; Li, W.; Zhu, Z.; Zhou, J.; and Lu, J. 2023 · 2023
Closest in time.
Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation
Stypułkowski, M.; Vougioukas, K.; He, S.; Zięba, M.; Petridis, S.; and Pantic, M. 2023 · 2023
Closest in time.
CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior
Xing, J.; Xia, M.; Zhang, Y.; Cun, X.; Wang, J.; and Wong, T.-T. 2023 · 2023
Closest in time.
ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model
Zhang, M.; Guo, X.; Pan, L.; Cai, Z.; Hong, F.; Li, H.; Yang, L.; and Liu, Z. 2023 · 2023
Closest in time.
Action2motion: Conditioned generation of 3d human motions
Guo, C.; Zuo, X.; Wang, S.; Zou, S.; Sun, Q.; Deng, A.; Gong, M.; and Cheng, L. 2020 · 2029
Closest in time.
Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory
Park, S. J.; Kim, M.; Hong, J.; Choi, J.; and Ro, Y. M. 2022 · 2070
Closest in time.