Fetching the paper…
Reading the bibliography…
Speech-driven facial animation methods usually contain two main classes, 3D and 2D talking face, both of which attract considerable research attention in recent years.
V. Blanz and T. Vetter, “A morphable model for the synthesis of 3d faces,” in Proc. of SIGGRAPH , 1999, pp. 187–194
1999
Earlier work this paper cites.
D. Vlasic, M. Brand, H. Pfister, and J. Popovic, “Face transfer with multilinear models,” ACM Trans. Graph. , vol. 24, no. 3, pp. 426–433, 2005
2005
Earlier work this paper cites.
G. Fanelli, J. Gall, H. Romsdorfer, T. Weise, and L. V. Gool, “A 3-d audio-visual corpus of affective communication,” IEEE Trans. Multim. , vol. 12, no. 6, pp. 591–598, 2010
2010
Earlier work this paper cites.
N. D. Narvekar and L. J. Karam, “A no-reference image blur metric based on the cumulative probability of blur detection (CPBD),” IEEE Trans. Image Process. , vol. 20, no. 9, pp. 2678–2683, 2011
2011
Earlier work this paper cites.
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio, “Generative adversarial nets,” in Proc. of NIPS , Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger, Eds., 2014, pp. 2672–2680
2014
Earlier work this paper cites.
C. Cao, Q. Hou, and K. Zhou, “Displaced dynamic expression regression for real-time facial tracking and animation,” ACM Trans. Graph. , vol. 33, no. 4, pp. 43:1–43:10, 2014
2014
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in Proc. of ICLR , Y. Bengio and Y. LeCun, Eds., 2014
2014
Earlier work this paper cites.
C. Cao, D. Bradley, K. Zhou, and T. Beeler, “Real-time high-fidelity facial performance capture,” ACM Trans. Graph. , vol. 34, no. 4, pp. 46:1–46:9, 2015
2015
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Out of time: Automated lip sync in the wild,” in Proc. of ACCV Workshop , C. Chen, J. Lu, and K. Ma, Eds., vol. 10117, 2016, pp. 251–263
2016
Earlier work this paper cites.
C. Wang, F. Shi, S. Xia, and J. Chai, “Realtime 3d eye gaze animation using a single RGB camera,” ACM Trans. Graph. , vol. 35, no. 4, pp. 118:1–118:14, 2016
2016
Earlier work this paper cites.
P. Garrido, M. Zollhöfer, D. Casas, L. Valgaerts, K. Varanasi, P. Pérez, and C. Theobalt, “Reconstruction of personalized 3d face rigs from monocular video,” ACM Trans. Graph. , vol. 35, no. 3, pp. 28:1–28:15, 2016
2016
Earlier work this paper cites.
J. S. Chung, A. Jamaludin, and A. Zisserman, “You said that?” in Proc. of BMVC , 2017
2017
Earlier work this paper cites.
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen, “Audio-driven facial animation by joint end-to-end learning of pose and emotion,” ACM Trans. Graph. , vol. 36, no. 4, pp. 94:1–94:12, 2017
2017
Earlier work this paper cites.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Proc. of NIPS , I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, Eds., 2017, pp. 6306–6315
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. of NIPS , I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, Eds., 2017, pp. 5998–6008
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in Proc. of NIPS , I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, Eds., 2017, pp. 6626–6637
2017
Earlier work this paper cites.
T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero, “Learning a model of facial shape and expression from 4d scans,” ACM Trans. Graph. , vol. 36, no. 6, pp. 194:1–194:17, 2017
2017
Earlier work this paper cites.
L. Chen, Z. Li, R. K. Maddox, Z. Duan, and C. Xu, “Lip movements generation at a glance,” in Proc. of ECCV , V. Ferrari, M. Hebert, C. Sminchisescu, and Y. Weiss, Eds., vol. 11211, 2018, pp. 538–553
2018
Earlier work this paper cites.
Y. Zhou, Z. Xu, C. Landreth, E. Kalogerakis, S. Maji, and K. Singh, “Visemenet: audio-driven animator-centric speech animation,” ACM Trans. Graph. , vol. 37, no. 4, p. 161, 2018
2018
Earlier work this paper cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in Proc. of CVPR , 2019, pp. 4401–4410
2019
Earlier work this paper cites.
K. R. Prajwal, R. Mukhopadhyay, J. Philip, A. Jha, V. P. Namboodiri, and C. V. Jawahar, “Towards automatic face-to-face translation,” in Proc. of ACM MM , L. Amsaleg, B. Huet, M. A. Larson, G. Gravier, H. Hung, C. Ngo, and W. T. Ooi, Eds., 2019, pp. 1428–1436
2019
Earlier work this paper cites.
L. Chen, R. K. Maddox, Z. Duan, and C. Xu, “Hierarchical cross-modal talking face generation with dynamic pixel-wise loss,” in Proc. of CVPR , 2019, pp. 7832–7841
2019
Earlier work this paper cites.
D. Cudeiro, T. Bolkart, C. Laidlaw, A. Ranjan, and M. J. Black, “Capture, learning, and synthesis of 3d speaking styles,” in Proc. of CVPR , 2019, pp. 10 101–10 111
2019
Earlier work this paper cites.
S. Sanyal, T. Bolkart, H. Feng, and M. J. Black, “Learning to regress 3d face shape and expression from an image without 3d supervision,” in Proc. of CVPR , 2019, pp. 7763–7772
2019
Earlier work this paper cites.
Y. Deng, J. Yang, S. Xu, D. Chen, Y. Jia, and X. Tong, “Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set,” in Proc. of CVPR Workshops , 2019, pp. 285–295
2019
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Proc. of NIPS , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., 2020
2020
Cited alongside, same era.
K. R. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. V. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in Proc. of ACM MM , C. W. Chen, R. Cucchiara, X. Hua, G. Qi, E. Ricci, Z. Zhang, and R. Zimmermann, Eds., 2020, pp. 484–492
2020
Cited alongside, same era.
K. Vougioukas, S. Petridis, and M. Pantic, “Realistic speech-driven facial animation with gans,” Int. J. Comput. Vis. , vol. 128, no. 5, pp. 1398–1413, 2020
2020
Cited alongside, same era.
P. Ma, S. Petridis, and M. Pantic, “Visual speech recognition for multiple languages in the wild,” Nat. Mac. Intell. , vol. 4, no. 11, pp. 930–939, 2022
2022
Later among the works it cites.
J. Liu, B. Hui, K. Li, Y. Liu, Y. Lai, Y. Zhang, Y. Liu, and J. Yang, “Geometry-guided dense perspective network for speech-driven facial animation,” IEEE Trans. Vis. Comput. Graph. , vol. 28, no. 12, pp. 4873–4886, 2022
2022
Later among the works it cites.
R. Danecek, M. J. Black, and T. Bolkart, “EMOCA: emotion driven monocular face capture and animation,” in Proc. of CVPR , 2022, pp. 20 279–20 290
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Chen, G. Cui, C. Liu, Z. Li, Z. Kou, Y. Xu, and C. Xu, “Talking-head generation with rhythmic head motion,” in Proc. of ECCV , A. Vedaldi, H. Bischof, T. Brox, and J. Frahm, Eds., vol. 12354, 2020, pp. 35–51
2020
Cited alongside, same era.
Y. Zhou, X. Han, E. Shechtman, J. Echevarria, E. Kalogerakis, and D. Li, “Makelttalk: speaker-aware talking-head animation,” ACM Trans. Graph. , vol. 39, no. 6, pp. 221:1–221:15, 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. of NIPS , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., 2020
2020
Cited alongside, same era.
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. of Interspeech , H. Meng, B. Xu, and T. F. Zheng, Eds., 2020, pp. 5036–5040
2020
Cited alongside, same era.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” in Proc. of ECCV , A. Vedaldi, H. Bischof, T. Brox, and J. Frahm, Eds., vol. 12346, 2020, pp. 405–421
2020
Cited alongside, same era.
H. Zhou, Y. Sun, W. Wu, C. C. Loy, X. Wang, and Z. Liu, “Pose-controllable talking face generation by implicitly modularized audio-visual representation,” in Proc. of CVPR , 2021, pp. 4176–4186
2021
Cited alongside, same era.
Y. Lu, J. Chai, and X. Cao, “Live speech portraits: real-time photorealistic talking-head animation,” ACM Trans. Graph. , vol. 40, no. 6, pp. 220:1–220:17, 2021
2021
Cited alongside, same era.
Z. Zhang, L. Li, Y. Ding, and C. Fan, “Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,” in Proc. of CVPR , 2021, pp. 3661–3670
2021
Cited alongside, same era.
S. Wang, L. Li, Y. Ding, and X. Yu, “One-shot talking face generation from single-speaker audio-visual correlation learning,” in Proc. of AAAI , 2022, pp. 2531–2539
2022
Later among the works it cites.
B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering,” ACM Trans. Graph. , vol. 42, no. 4, pp. 139:1–139:14, 2023
2023
Later among the works it cites.
R. Liu, C. Li, H. Cao, Y. Zheng, M. Zeng, and X. Cheng, “EMEF: ensemble multi-exposure image fusion,” in Proc. of AAAI , 2023, pp. 1710–1718
2023
Later among the works it cites.
W. Zhang, X. Cun, X. Wang, Y. Zhang, X. Shen, Y. Guo, Y. Shan, and F. Wang, “Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,” in Proc. of CVPR , 2023, pp. 8652–8661
2023
Later among the works it cites.
Z. Zhang, Z. Hu, W. Deng, C. Fan, T. Lv, and Y. Ding, “Dinet: Deformation inpainting network for realistic face visually dubbing on high resolution video,” in Proc. of AAAI , B. Williams, Y. Chen, and J. Neville, Eds., 2023, pp. 3543–3551
2023
Later among the works it cites.
J. Xing, M. Xia, Y. Zhang, X. Cun, J. Wang, and T. Wong, “Codetalker: Speech-driven 3d facial animation with discrete motion prior,” in Proc. of CVPR , 2023, pp. 12 780–12 790
2023
Later among the works it cites.
S. Stan, K. I. Haque, and Z. Yumak, “Facediffuser: Speech-driven 3d facial animation synthesis using diffusion,” in ACM Conference on Motion, Interaction and Games , J. Pettré, B. Solenthaler, R. McDonnell, and C. Peters, Eds., 2023, pp. 13:1–13:11
2023
Later among the works it cites.
H. Yi, H. Liang, Y. Liu, Q. Cao, Y. Wen, T. Bolkart, D. Tao, and M. J. Black, “Generating holistic 3d human motion from speech,” in Proc. of CVPR , 2023, pp. 469–480
2023
Later among the works it cites.
S. Alexanderson, R. Nagy, J. Beskow, and G. E. Henter, “Listen, denoise, action! audio-driven motion synthesis with diffusion models,” ACM Trans. Graph. , vol. 42, no. 4, pp. 44:1–44:20, 2023
2023
Later among the works it cites.
Z. Zhou and B. Wang, “UDE: A unified driving engine for human motion generation,” in Proc. of CVPR , 2023, pp. 5632–5641
2023
Later among the works it cites.
J. Lin, J. Chang, L. Liu, G. Li, L. Lin, Q. Tian, and C. W. Chen, “Being comes from not-being: Open-vocabulary text-to-motion generation with wordless training,” in Proc. of CVPR , 2023, pp. 23 222–23 231
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Liu, Y. Cheng, S. Huang, C. Li, and X. Cheng, “Transformer-based high-fidelity facial displacement completion for detailed 3d face reconstruction,” IEEE Trans. on Multim. , pp. 1–13, 2023
2023
Later among the works it cites.
P. Ma, A. Haliassos, A. Fernandez-Lopez, H. Chen, S. Petridis, and M. Pantic, “Auto-avsr: Audio-visual speech recognition with automatic labels,” in Proc. of ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.