Fetching the paper…
Reading the bibliography…
Human-human communication is like a delicate dance where listeners and speakers concurrently interact to maintain conversational dynamics.
Cao, Y., Tien, W.C., Faloutsos, P., Pighin, F.: Expressive speech-driven facial animation. ACM Transactions on Graphics 24
2005
Earlier work this paper cites.
Gratch, J., Wang, N., Gerten, J., Fast, E., Duffy, R.: Creating rapport with virtual agents. In: Intelligent Virtual Agents (IVA). pp. 125–138. Springer Berlin Heidelberg, Berlin, Heidelberg (2007)
2007
Earlier work this paper cites.
Messinger, D.S., Mahoor, M.H., Chow, S.M., Cohn, J.F.: Automated measurement of facial expression in infant–mother interaction: A pilot study. Infancy 14
2009
Earlier work this paper cites.
Bohus, D., Horvitz, E.: Facilitating multiparty dialog with gaze, gesture, and speech. In: International Conference on Multimodal Interfaces and the Workshop on Machine Learning for Multimodal Interaction. pp. 1–8 (2010)
2010
Earlier work this paper cites.
Fanelli, G., Gall, J., Romsdorfer, H., Weise, T., Van Gool, L.: A 3-d audio-visual corpus of affective communication. IEEE Transactions on Multimedia 12
2010
Earlier work this paper cites.
Weise, T., Bouaziz, S., Li, H., Pauly, M.: Realtime performance-based facial animation. ACM Transactions on Graphics 30
2011
Earlier work this paper cites.
Massaro, D., Cohen, M., Tabain, M., Beskow, J., Clark, R.: Animated speech: research progress and applications. Audiovisual Speech Processing p. 309–345 (2012)
2012
Earlier work this paper cites.
Taylor, S.L., Mahler, M., Theobald, B.J., Matthews, I.: Dynamic units of visual speech. In: Proceedings of the ACM SIGGRAPH/Eurographics conference on Computer Animation. pp. 275–284 (2012)
2012
Earlier work this paper cites.
DeVito, J.A.: Interpersonal communication book, The, 13/E. Pearson, London, UK (2013)
2013
Earlier work this paper cites.
Li, H., Yu, J., Ye, Y., Bregler, C.: Realtime facial animation with on-the-fly correctives. ACM Transactions on Graphics 32
2013
Earlier work this paper cites.
Xu, Y., Feng, A.W., Marsella, S., Shapiro, A.: A practical and configurable lip sync method for games. In: Proceedings of Motion on Games. pp. 131–140 (2013)
2013
Earlier work this paper cites.
Fan, B., Wang, L., Soong, F.K., Xie, L.: Photo-real talking head with deep bidirectional lstm. In: IEEE International Conference on Acoustics, Speech and Signal Processing. pp. 4884–4888. IEEE (2015)
2015
Earlier work this paper cites.
Liu, Y., Xu, F., Chai, J., Tong, X., Wang, L., Huo, Q.: Video-audio driven real-time facial animation. ACM Transactions on Graphics 34
2015
Earlier work this paper cites.
Cao, C., Wu, H., Weng, Y., Shao, T., Zhou, K.: Real-time facial animation with image-based dynamic avatars. ACM Transactions on Graphics 35
2016
Earlier work this paper cites.
Cerekovic, A., Aran, O., Gatica-Perez, D.: Rapport with virtual agents: What do human social cues and personality explain? IEEE Transactions on Affective Computing 8
2016
Earlier work this paper cites.
Chung, J.S., Zisserman, A.: Out of time: automated lip sync in the wild. In: Asian Conference on Computer Vision. pp. 251–263. Springer (2016)
2016
Earlier work this paper cites.
Edwards, P., Landreth, C., Fiume, E., Singh, K.: Jali: an animator-centric viseme model for expressive lip synchronization. ACM Transactions on Graphics) 35
2016
Earlier work this paper cites.
Greenwood, D., Laycock, S., Matthews, I.: Predicting head pose in dyadic conversation. In: International Conference on Intelligent Virtual Agents. pp. 160–169. Springer (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Karras, T., Aila, T., Laine, S., Herva, A., Lehtinen, J.: Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Transactions on Graphics 36
2017
Earlier work this paper cites.
Li, T., Bolkart, T., Black, M.J., Li, H., Romero, J.: Learning a model of facial shape and expression from 4D scans. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia) 36
2017
Earlier work this paper cites.
Mirsamadi, S., Barsoum, E., Zhang, C.: Automatic speech emotion recognition using recurrent neural networks with local attention. In: 2017 IEEE International conference on acoustics, speech and signal processing (ICASSP). pp. 2227–2231. IEEE (2017)
2017
Earlier work this paper cites.
van den Oord, A., Vinyals, O., Kavukcuoglu, K.: Neural discrete representation learning. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. pp. 6309–6318 (2017)
2017
Earlier work this paper cites.
Riehle, M., Kempkensteffen, J., Lincoln, T.M.: Quantifying facial expression synchrony in face-to-face dyadic interactions: Temporal dynamics of simultaneously recorded facial emg signals. Journal of Nonverbal Behavior 41
2017
Earlier work this paper cites.
Suwajanakorn, S., Seitz, S.M., Kemelmacher-Shlizerman, I.: Synthesizing obama: learning lip sync from audio. ACM Transactions on Graphics 36
2017
Earlier work this paper cites.
Taylor, S., Kim, T., Yue, Y., Mahler, M., Krahe, J., Rodriguez, A.G., Hodgins, J., Matthews, I.: A deep learning approach for generalized speech animation. ACM Transactions on Graphics 36
2017
Earlier work this paper cites.
Yu, J., Chen, C.W.: From talking head to singing head: a significant enhancement for more natural human computer interaction. In: 2017 IEEE International Conference on Multimedia and Expo (ICME). pp. 511–516. IEEE (2017)
2017
Earlier work this paper cites.
Chen, L., Li, Z., Maddox, R.K., Duan, Z., Xu, C.: Lip movements generation at a glance. In: Proceedings of the European Conference on Computer Vision. pp. 520–535 (2018)
2018
Earlier work this paper cites.
Chu, H., Li, D., Fidler, S.: A face-to-face neural conversation model. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 7113–7121 (2018)
2018
Earlier work this paper cites.
Kim, H., Garrido, P., Tewari, A., Xu, W., Thies, J., Niessner, M., Pérez, P., Richardt, C., Zollhöfer, M., Theobalt, C.: Deep video portraits. ACM Transactions on Graphics 37
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Pham, H.X., Wang, Y., Pavlovic, V.: End-to-end learning for 3d facial animation from speech. In: Proceedings of the ACM International Conference on Multimodal Interaction. pp. 361–365 (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Li, R., Yang, S., Ross, D.A., Kanazawa, A.: Ai choreographer: Music conditioned 3d dance generation with aist++. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 13401–13412 (2021)
2021
Later among the works it cites.
Ren, Y., Li, G., Chen, Y., Li, T.H., Liu, S.: Pirenderer: Controllable portrait image generation via semantic neural rendering. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 13759–13768 (2021)
2021
Later among the works it cites.
Richard, A., Zollhöfer, M., Wen, Y., de la Torre, F., Sheikh, Y.: Meshtalk: 3d face animation from speech using cross-modality disentanglement. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 1173–1182 (2021)
2021
Later among the works it cites.
Song, L., Liu, B., Yin, G., Dong, X., Zhang, Y., Bai, J.X.: Tacr-net: Editing on deep video and voice portraits. In: Proceedings of the 29th ACM International Conference on Multimedia. pp. 478–486 (2021)
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Zhou, Y., Xu, Z., Landreth, C., Kalogerakis, E., Maji, S., Singh, K.: Visemenet: Audio-driven animator-centric speech animation. ACM Transactions on Graphics 37
2018
Cited alongside, same era.
Zollhöfer, M., Thies, J., Garrido, P., Bradley, D., Beeler, T., Pérez, P., Stamminger, M., Nießner, M., Theobalt, C.: State of the art on monocular 3d face reconstruction, tracking, and applications. In: Computer Graphics Forum. pp. 523–550 (2018)
2018
Cited alongside, same era.
Ahuja, C., Ma, S., Morency, L.P., Sheikh, Y.: To react or not to react: End-to-end visual pose forecasting for personalized avatar during dyadic conversations. In: 2019 International conference on multimodal interaction. pp. 74–84 (2019)
2019
Cited alongside, same era.
Chen, L., Maddox, R.K., Duan, Z., Xu, C.: Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 7832–7841 (2019)
2019
Cited alongside, same era.
Cudeiro, D., Bolkart, T., Laidlaw, C., Ranjan, A., Black, M.J.: Capture, learning, and synthesis of 3d speaking styles. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 10101–10111 (2019)
2019
Cited alongside, same era.
Fried, O., Tewari, A., Zollhöfer, M., Finkelstein, A., Shechtman, E., Goldman, D.B., Genova, K., Jin, Z., Theobalt, C., Agrawala, M.: Text-based editing of talking-head video. ACM Transactions on Graphics 38
2019
Cited alongside, same era.
Jonell, P., Kucherenko, T., Ekstedt, E., Beskow, J.: Learning non-verbal behavior for a social robot from youtube videos. In: ICDL-EpiRob Workshop on Naturalistic Non-Verbal and Affective Human-Robot Interactions, Oslo, Norway, August 19, 2019 (2019)
2019
Cited alongside, same era.
Later among the works it cites.
Song, L., Liu, B., Yu, N.: Talking face video generation with editable expression. In: Image and Graphics: 11th International Conference, ICIG 2021, Haikou, China, August 6–8, 2021, Proceedings, Part III 11. pp. 753–764. Springer (2021)
2021
Later among the works it cites.
Song, L., Yin, G., Liu, B., Zhang, Y., Yu, N.: Fsft-net: face transfer video generation with few-shot views. In: 2021 IEEE International Conference on Image Processing (ICIP). pp. 3582–3586. IEEE (2021)
2021
Later among the works it cites.
Zhang, Z., Li, L., Ding, Y., Fan, C.: Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3661–3670 (2021)
2021
Later among the works it cites.
Zhou, H., Sun, Y., Wu, W., Loy, C.C., Wang, X., Liu, Z.: Pose-controllable talking face generation by implicitly modularized audio-visual representation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 4176–4186 (2021)
2021
Later among the works it cites.
Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T. (eds.): Computer Vision – ECCV 2022. Springer (2022). https://doi.org/10.1007/978-3-031-19769-7
2022
Later among the works it cites.
Danecek, R., Black, M.J., Bolkart, T.: EMOCA: Emotion driven monocular face capture and animation. In: Conference on Computer Vision and Pattern Recognition (CVPR). pp. 20311–20322 (2022)
2022
Later among the works it cites.
Fan, Y., Lin, Z., Saito, J., Wang, W., Komura, T.: Faceformer: Speech-driven 3d facial animation with transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 18770–18780 (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Ng, E., Joo, H., Hu, L., Li, H., Darrell, T., Kanazawa, A., Ginosar, S.: Learning to listen: Modeling non-deterministic dyadic facial motion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 20395–20405 (2022)
2022
Later among the works it cites.
Palmero, C., Barquero, G., Junior, J.C.J., Clapés, A., Núnez, J., Curto, D., Smeureanu, S., Selva, J., Zhang, Z., Saeteros, D., et al.: Chalearn lap challenges on self-reported personality recognition and non-verbal behavior forecasting during social dyadic interactions: Dataset, design, and results. In: Understanding Social Behavior in Dyadic and Small Group Interactions. pp. 4–52. PMLR (2022)
2022
Later among the works it cites.
Song, L., Fang, Z., Li, X., Dong, X., Jin, Z., Chen, Y., Lyu, S.: Adaptive Face Forgery Detection in Cross Domain. In: Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T. (eds.) Computer Vision – ECCV 2022. pp. 467–484. Lecture Notes in Computer Science, Springer Nature Switzerland, Cham (2022). https://doi.org/10.1007/978-3-031-19830-4_27
2022
Later among the works it cites.
Song, L., Li, X., Fang, Z., Jin, Z., Chen, Y., Xu, C.: Face forgery detection via symmetric transformer. In: Proceedings of the 30th ACM International Conference on Multimedia. pp. 4102–4111 (2022)
2022
Later among the works it cites.
Zhou, M., Bai, Y., Zhang, W., Yao, T., Zhao, T., Mei, T.: Responsive listening head generation: a benchmark dataset and baseline. In: European Conference on Computer Vision. pp. 124–142. Springer (2022)
2022
Later among the works it cites.
Chang, Z., Hu, W., Yang, Q., Zheng, S.: Hierarchical semantic perceptual listener head video generation: A high-performance pipeline. In: Proceedings of the 31st ACM International Conference on Multimedia. pp. 9581–9585 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Kucherenko, T., Nagy, R., Yoon, Y., Woo, J., Nikolov, T., Tsakov, M., Henter, G.E.: The genea challenge 2023: A large-scale evaluation of gesture generation models in monadic and dyadic settings. In: Proceedings of the 25th International Conference on Multimodal Interaction. pp. 792–801 (2023)
2023
Later among the works it cites.
Ng, E., Subramanian, S., Klein, D., Kanazawa, A., Darrell, T., Ginosar, S.: Can language models learn to listen? In: Proceedings of the International Conference on Computer Vision (ICCV) (2023)
2023
Later among the works it cites.
Reece, A., Cooney, G., Bull, P., Chung, C., Dawson, B., Fitzpatrick, C., Glazer, T., Knox, D., Liebscher, A., Marin, S.: The candor corpus: Insights from a large multimodal dataset of naturalistic conversation. Science Advances 9
2023
Later among the works it cites.
Song, L., Yin, G., Jin, Z., Dong, X., Xu, C.: Emotional listener portrait: Realistic listener motion simulation in conversation. In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV). pp. 20782–20792. IEEE (2023)
2023
Later among the works it cites.
Song, S., Spitale, M., Luo, C., Barquero, G., Palmero, C., Escalera, S., Valstar, M., Baur, T., Ringeval, F., André, E., et al.: React2023: The first multiple appropriate facial reaction generation challenge. In: Proceedings of the 31st ACM International Conference on Multimedia. pp. 9620–9624 (2023)
2023
Later among the works it cites.
Stan, S., Haque, K.I., Yumak, Z.: Facediffuser: Speech-driven 3d facial animation synthesis using diffusion. In: Proceedings of the 16th ACM SIGGRAPH Conference on Motion, Interaction and Games. pp. 1–11 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Xing, J., Xia, M., Zhang, Y., Cun, X., Wang, J., Wong, T.T.: Codetalker: Speech-driven 3d facial animation with discrete motion prior. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 12780–12790 (2023)
2023
Later among the works it cites.