Fetching the paper…
Reading the bibliography…
We propose DiffSHEG, a Diffusion-based approach for Speech-driven Holistic 3D Expression and Gesture generation with arbitrary length.
The persona effect: How substantial is it?
Susanne Van Mulken, Elisabeth André, and Jochen Müller · 1998
Earlier work this paper cites.
The role of gesture in communication and thinking
Susan Goldin-Meadow · 1999
Earlier work this paper cites.
Beat: the behavior expression animation toolkit
Justine Cassell, Hannes Högni Vilhjálmsson, and Timothy Bickmore · 2004
Earlier work this paper cites.
Gesture Generation by Imitation: From Human Behavior to Computer Character Animation
Michael Kipp · 2004
Earlier work this paper cites.
Expressive speech-driven facial animation
Yong Cao, Wen C Tien, Petros Faloutsos, and Frédéric Pighin · 2005
Earlier work this paper cites.
Towards a common framework for multimodal generation: The behavior markup language
Stefan Kopp, Brigitte Krenn, Stacy Marsella, Andrew N Marshall, Catherine Pelachaud, Hannes Pirker, Kristinn R Thórisson, and Hannes Vilhjálmsson · 2006
Earlier work this paper cites.
Robot behavior toolkit: generating effective social behaviors for robots
Chien-Ming Huang and Bilge Mutlu · 2012
Earlier work this paper cites.
Animated speech: research progress and applications
DW Massaro, MM Cohen, M Tabain, J Beskow, and R Clark · 2012
Earlier work this paper cites.
Dynamic units of visual speech
Sarah L Taylor, Moshe Mahler, Barry-John Theobald, and Iain Matthews · 2012
Earlier work this paper cites.
A practical and configurable lip sync method for games
Yuyu Xu, Andrew W Feng, Stacy Marsella, and Ari Shapiro · 2013
Earlier work this paper cites.
Gesture and speech in interaction: An overview, 2014
Petra Wagner, Zofia Malisz, and Stefan Kopp · 2014
Earlier work this paper cites.
Video-audio driven real-time facial animation
Yilong Liu, Feng Xu, Jinxiang Chai, Xin Tong, Lijuan Wang, and Qiang Huo · 2015
Earlier work this paper cites.
Jali: an animator-centric viseme model for expressive lip synchronization
Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh · 2016
Earlier work this paper cites.
Vid2speech: Speech reconstruction from silent video
Ariel Ephrat and Shmuel Peleg · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization
Xun Huang and Serge J. Belongie · 2017
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Earlier work this paper cites.
Emotion recognition using facial expressions
Paweł Tarnowski, Marcin Kołodziej, Andrzej Majkowski, and Remigiusz J. Rak · 2017
Earlier work this paper cites.
A deep learning approach for generalized speech animation
Sarah Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews · 2017
Cited alongside, same era.
Evaluation of speech-to-gesture generation using bi-directional lstm network
Dai Hasegawa, Naoshi Kaneko, Shinichi Shirakawa, Hiroshi Sakuta, and Kazuhiko Sumi · 2018
Cited alongside, same era.
End-to-end learning for 3d facial animation from speech
Hai Xuan Pham, Yuting Wang, and Vladimir Pavlovic · 2018
Cited alongside, same era.
Naoqi api documentation
Robotics Softbank · 2018
Cited alongside, same era.
Capture, learning, and synthesis of 3d speaking styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black · 2019
Cited alongside, same era.
Expressive body capture: 3D hands, face, and body from a single image
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black · 2019
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2021
Later among the works it cites.
Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings
Tenglong Ao, Qingzhe Gao, Yuke Lou, Baoquan Chen, and Libin Liu · 2022
Later among the works it cites.
Faceformer: Speech-driven 3d facial animation with transformers
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura · 2022
Later among the works it cites.
Repaint: Inpainting using denoising diffusion probabilistic models
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool · 2022
Later among the works it cites.
Real-time streaming video denoising with bidirectional buffers
Chenyang Qi, Junming Chen, Xin Yang, and Qifeng Chen · 2022
Later among the works it cites.
Audio-driven stylized gesture generation with flow-based model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots
Youngwoo Yoon, Woo-Ri Ko, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Modality dropout for improved performance-driven talking faces
Ahmed Hussen Abdelaziz, Barry-John Theobald, Paul Dixon, Reinhard Knothe, Nicholas Apostoloff, and Sachin Kajareker · 2020
Cited alongside, same era.
Gesticulator: A framework for semantically-aware speech-driven gesture generation
Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav Eje Henter, Simon Alexandersson, Iolanda Leite, and Hedvig Kjellström · 2020
Cited alongside, same era.
Speech gesture generation from the trimodal context of text, audio, and speaker identity
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee · 2020
Cited alongside, same era.
Sheng Ye, Yu-Hui Wen, Yanan Sun, Ying He, Ziyang Zhang, Yaoyuan Wang, Weihua He, and Yong-Jin Liu · 2022
Later among the works it cites.
Listen, denoise, action! audio-driven motion synthesis with diffusion models
Simon Alexanderson, Rajmund Nagy, Jonas Beskow, and Gustav Eje Henter · 2023
Later among the works it cites.
Gesturediffuclip: Gesture diffusion model with clip latents
Tenglong Ao, Zeyi Zhang, and Libin Liu · 2023
Later among the works it cites.
Moda: Mapping-once audio-driven portrait animation with dual attentions
Yunfei Liu, Lijian Lin, Fei Yu, Changyin Zhou, and Yu Li · 2023
Later among the works it cites.
Difftalk: Crafting diffusion models for generalized talking head synthesis
Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu · 2023
Later among the works it cites.
Codetalker: Speech-driven 3d facial animation with discrete motion prior
Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong · 2023
Later among the works it cites.
Diffusestylegesture: Stylized audio-driven co-speech gesture generation with diffusion models
Sicheng Yang, Zhiyong Wu, Minglei Li, Zhensong Zhang, Lei Hao, Weihong Bao, Ming Cheng, and Long Xiao · 2023
Later among the works it cites.
Generating holistic 3d human motion from speech
Hongwei Yi, Hualin Liang, Yifei Liu, Qiong Cao, Yandong Wen, Timo Bolkart, Dacheng Tao, and Michael J Black · 2023
Later among the works it cites.
Taming diffusion models for audio-driven co-speech gesture generation
Lingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian, Ziwei Liu, and Lequan Yu · 2023
Later among the works it cites.
https://www.unrealengine.com/en-US/metahuman , 2023
MetaHuman — Realistic Person Creator · 2023
Later among the works it cites.
https://www.unrealengine.com/en-US/unreal-engine-5 , 2023
Unreal Engine 5 · 2023
Later among the works it cites.
Speech2affectivegestures: Synthesizing co-speech gestures with generative adversarial affective expression learning
Uttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, and Dinesh Manocha · 2036
Closest in time.