Fetching the paper…
Reading the bibliography…
The creation of lifelike speech-driven 3D facial animation requires a natural and precise synchronization between audio input and facial expressions.
Fréchet, M.: Sur la distance de deux lois de probabilité. In: Annales de l’ISUP, vol. 6, pp. 183–198 (1957)
1957
Earlier work this paper cites.
Blanz, V., Vetter, T.: A morphable model for the synthesis of 3d faces. In: Proceedings of the 26th Annual Conference on Computer Graphics and Interactive Techniques, pp. 187–194 (1999)
1999
Earlier work this paper cites.
Ezzat, T., Poggio, T.: Visual speech synthesis by morphing visemes. International Journal of Computer Vision 38
2000
Earlier work this paper cites.
Cohen, M.M., Clark, R., Massaro, D.W.: Animated speech: Research progress and applications. In: AVSP 2001-International Conference on Auditory-Visual Speech Processing (2001)
2001
Earlier work this paper cites.
Fanelli, G., Gall, J., Romsdorfer, H., Weise, T., Van Gool, L.: A 3-d audio-visual corpus of affective communication. IEEE Transactions on Multimedia 12
2010
Earlier work this paper cites.
Taylor, S.L., Mahler, M., Theobald, B.-J., Matthews, I.: Dynamic units of visual speech. In: Proceedings of the 11th ACM SIGGRAPH/Eurographics Conference on Computer Animation, pp. 275–284 (2012)
2012
Earlier work this paper cites.
Xu, Y., Feng, A.W., Marsella, S., Shapiro, A.: A practical and configurable lip sync method for games. In: Proceedings of Motion on Games, pp. 131–140 (2013)
2013
Earlier work this paper cites.
Fan, B., Wang, L., Soong, F.K., Xie, L.: Photo-real talking head with deep bidirectional lstm. In: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4884–4888 (2015). IEEE
2015
Earlier work this paper cites.
Edwards, P., Landreth, C., Fiume, E., Singh, K.: Jali: an animator-centric viseme model for expressive lip synchronization. ACM Transactions on graphics (TOG) 35
2016
Earlier work this paper cites.
Van Den Oord, A., Vinyals, O., et al.: Neural discrete representation learning. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
Chung, J.S., Zisserman, A.: Out of time: automated lip sync in the wild. In: Computer Vision–ACCV 2016 Workshops: ACCV 2016 International Workshops, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part II 13, pp. 251–263 (2017). Springer
2017
Earlier work this paper cites.
Suwajanakorn, S., Seitz, S.M., Kemelmacher-Shlizerman, I.: Synthesizing obama: learning lip sync from audio. ACM Transactions on Graphics (ToG) 36
2017
Earlier work this paper cites.
Karras, T., Aila, T., Laine, S., Herva, A., Lehtinen, J.: Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Transactions on Graphics (TOG) 36
2017
Earlier work this paper cites.
Li, T., Bolkart, T., Black, M.J., Li, H., Romero, J.: Learning a model of facial shape and expression from 4d scans. ACM Trans. Graph. 36
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Chen, L., Li, Z., Maddox, R.K., Duan, Z., Xu, C.: Lip movements generation at a glance. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 520–535 (2018)
2018
Earlier work this paper cites.
Zhou, Y., Xu, Z., Landreth, C., Kalogerakis, E., Maji, S., Singh, K.: Visemenet: Audio-driven animator-centric speech animation. ACM Transactions on Graphics (TOG) 37
2018
Earlier work this paper cites.
Cudeiro, D., Bolkart, T., Laidlaw, C., Ranjan, A., Black, M.J.: Capture, learning, and synthesis of 3d speaking styles. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10101–10111 (2019)
2019
Cited alongside, same era.
Zhang, J., Fisher, R.B.: 3d visual passcode: Speech-driven 3d facial dynamics for behaviometrics. Signal processing 160
2019
Cited alongside, same era.
Chiu, H.-k., Adeli, E., Wang, B., Huang, D.-A., Niebles, J.C.: Action-agnostic human pose forecasting. In: 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 1423–1432 (2019). IEEE
2019
Cited alongside, same era.
Du, X., Vasudevan, R., Johnson-Roberson, M.: Bio-lstm: A biomechanically inspired recurrent neural network for 3-d pedestrian pose and gait prediction. IEEE Robotics and Automation Letters 4
2019
Cited alongside, same era.
Liu, J., Hui, B., Li, K., Liu, Y., Lai, Y.-K., Zhang, Y., Liu, Y., Yang, J.: Geometry-guided dense perspective network for speech-driven facial animation. IEEE Transactions on Visualization and Computer Graphics 28
2021
Later among the works it cites.
Doukas, M.C., Zafeiriou, S., Sharmanska, V.: Headgan: One-shot neural head synthesis and editing. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 14398–14407 (2021)
2021
Later among the works it cites.
Zhou, H., Sun, Y., Wu, W., Loy, C.C., Wang, X., Liu, Z.: Pose-controllable talking face generation by implicitly modularized audio-visual representation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4176–4186 (2021)
2021
Later among the works it cites.
Lu, Y., Chai, J., Cao, X.: Live speech portraits: real-time photorealistic talking-head animation. ACM Transactions on Graphics (TOG) 40
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, L., Maddox, R.K., Duan, Z., Xu, C.: Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7832–7841 (2019)
2019
Cited alongside, same era.
Ackland, S., Chiclana, F., Istance, H., Coupland, S.: Real-time 3d head pose tracking through 2.5 d constrained local models with local neural fields. International Journal of Computer Vision 127
2019
Cited alongside, same era.
Chen, L., Cui, G., Liu, C., Li, Z., Kou, Z., Xu, Y., Xu, C.: Talking-head generation with rhythmic head motion. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IX, pp. 35–51 (2020). Springer
2020
Cited alongside, same era.
Das, D., Biswas, S., Sinha, S., Bhowmick, B.: Speech-driven facial animation using cascaded gans for learning of motion and texture. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXX 16, pp. 408–424 (2020). Springer
2020
Cited alongside, same era.
Prajwal, K., Mukhopadhyay, R., Namboodiri, V.P., Jawahar, C.: A lip sync expert is all you need for speech to lip generation in the wild. In: Proceedings of the 28th ACM International Conference on Multimedia, pp. 484–492 (2020)
2020
Cited alongside, same era.
Vougioukas, K., Petridis, S., Pantic, M.: Realistic speech-driven facial animation with gans. International Journal of Computer Vision 128
2020
Cited alongside, same era.
Zhou, Y., Han, X., Shechtman, E., Echevarria, J., Kalogerakis, E., Li, D.: Makelttalk: speaker-aware talking-head animation. ACM Transactions On Graphics (TOG) 39
2020
Cited alongside, same era.
Pumarola, A., Agudo, A., Martinez, A.M., Sanfeliu, A., Moreno-Noguer, F.: Ganimation: One-shot anatomically consistent facial animation. International Journal of Computer Vision 128
2020
Cited alongside, same era.
Zhang, C., Zhao, Y., Huang, Y., Zeng, M., Ni, S., Budagavi, M., Guo, X.: Facial: Synthesizing dynamic talking face with implicit attribute learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3867–3876 (2021)
2021
Later among the works it cites.
Fan, Y., Lin, Z., Saito, J., Wang, W., Komura, T.: Faceformer: Speech-driven 3d facial animation with transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18770–18780 (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Yin, F., Zhang, Y., Cun, X., Cao, M., Fan, Y., Wang, X., Bai, Q., Wu, B., Wang, J., Yang, Y.: Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan. In: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XVII, pp. 85–101 (2022). Springer
2022
Later among the works it cites.
Chen, Z., Huang, Y., Yu, H., Wang, L.: Learning a robust part-aware monocular 3d human pose estimator via neural architecture search. International Journal of Computer Vision, 1–20 (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
Kang, Z., Sadeghi, M., Horaud, R., Alameda-Pineda, X.: Expression-preserving face frontalization improves visually assisted speech processing. International Journal of Computer Vision 131
2023
Closest in time.
2023
Closest in time.
Garg, R., Gao, R., Grauman, K.: Visually-guided audio spatialization in video with geometry-aware multi-task learning. International Journal of Computer Vision, 1–15 (2023)
2023
Closest in time.