Fetching the paper…
Reading the bibliography…
We propose the first approach to automatically and jointly synthesize both the synchronous 3D conversational body and hand gestures, as well as 3D face and head animations, of a virtual character from speech input.
Animated conversation: Rule-based generation of facial expression, gesture & spoken intonation for multiple conversational agents
Justine Cassell, Catherine Pelachaud, Norman Badler, Mark Steedman, Brett Achorn, Tripp Becket, Brett Douville, Scott Prevost, and Matthew Stone · 1994
Earlier work this paper cites.
The persona effect: How substantial is it?
Susanne Van Mulken, Elisabeth André, and Jochen Müller · 1998
Earlier work this paper cites.
The role of gesture in communication and thinking
Susan Goldin-Meadow · 1999
Earlier work this paper cites.
Embodied conversational interface agents
Justine Cassell · 2000
Earlier work this paper cites.
Language and Gesture
David McNeill · 2000
Earlier work this paper cites.
The CMU SPHINX-4 speech recognition system, 2003
Paul Lamere, Philip Kwok, Evandro Gouvêa, Bhiksha Raj, Rita Singh, William Walker, Manfred Warmuth, and Peter Wolf · 2003
Earlier work this paper cites.
BEAT: the Behavior Expression Animation Toolkit
Justine Cassell, Hannes Högni Vilhjálmsson, and Timothy Bickmore · 2004
Earlier work this paper cites.
Gesture: Visible Action as Utterance
Adam Kendon · 2004
Earlier work this paper cites.
Gesture modeling and animation based on a probabilistic re-creation of speaker style
Michael Neff, Michael Kipp, Irene Albrecht, and Hans-Peter Seidel · 2008
Earlier work this paper cites.
Real-time prosody-driven synthesis of body language
Sergey Levine, Christian Theobalt, and Vladlen Koltun · 2009
Earlier work this paper cites.
Gesture controllers
Sergey Levine, Philipp Krähenbühl, Sebastian Thrun, and Vladlen Koltun · 2010
Earlier work this paper cites.
How to train your avatar: A data driven approach to gesture generation
Chung-Cheng Chiu and Stacy Marsella · 2011
Earlier work this paper cites.
Deformable model fitting by regularized landmark mean-shift
Jason M. Saragih, Simon Lucey, and Jeffrey F. Cohn · 2011
Earlier work this paper cites.
Generating human-like behaviors using joint, speech-driven models for conversational agents
S. Mariooryad and C. Busso · 2012
Earlier work this paper cites.
Multimodal analysis of speech prosody and upper body gestures using hidden semi-markov models
E. Bozkurt, S. Asta, S. Özkul, Y. Yemez, and E. Erzin · 2013
Earlier work this paper cites.
Virtual character performance from speech
Stacy Marsella, Yuyu Xu, Margaux Lhommet, Andrew Feng, Stefan Scherer, and Ari Shapiro · 2013
Earlier work this paper cites.
Gesture generation with low-dimensional embeddings
Chung-Cheng Chiu and Stacy Marsella · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition, 2014
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, and Andrew Y. Ng · 2014
Earlier work this paper cites.
Speech-driven animation constrained by appropriate discourse functions
Najmeh Sadoughi, Yang Liu, and Carlos Busso · 2014
Earlier work this paper cites.
Video-audio driven real-time facial animation
Yilong Liu, Feng Xu, Jinxiang Chai, Xin Tong, Lijuan Wang, and Qiang Huo · 2015
Cited alongside, same era.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Cited alongside, same era.
FFMPEG. ffmpeg.org, 2016
FFmpeg Developers · 2016
Cited alongside, same era.
Reconstruction of personalized 3d face rigs from monocular video
Pablo Garrido, Michael Zollhöfer, Dan Casas, Levi Valgaerts, Kiran Varanasi, Patrick Pérez, and Christian Theobalt · 2016
Cited alongside, same era.
Bidirectional lstm networks employing stacked bottleneck features for expressive speech-driven head motion synthesis
Kathrin Haag and Hiroshi Shimodaira · 2016
Cited alongside, same era.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, Xi Chen, and Xi Chen · 2016
Novel realizations of speech-driven head movements with generative adversarial networks
N. Sadoughi and C. Busso · 2018
Later among the works it cites.
Audio to body dynamics
Eli Shlizerman, Lucio Dery, Hayden Schoen, and Ira Kemelmacher-Shlizerman · 2018
Later among the works it cites.
github.com/usuyama/pytorch-unet, 2018
Naoto Usuyama · 2018
Later among the works it cites.
Capture, learning, and synthesis of 3D speaking styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael Black · 2019
Later among the works it cites.
Multi-objective adversarial gesture generation
Ylva Ferstl, Michael Neff, and Rachel McDonnell · 2019
Later among the works it cites.
Learning individual styles of conversational gesture
S. Ginosar, A. Bar, G. Kohavi, C. Chan, A. Owens, and J. Malik · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2017
Cited alongside, same era.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Cited alongside, same era.
Learning a model of facial shape and expression from 4D scans
Tianye Li, Timo Bolkart, Michael J. Black, Hao Li, and Javier Romero · 2017
Cited alongside, same era.
A simple yet effective baseline for 3d human pose estimation
Julieta Martinez, Rayat Hossain, Javier Romero, and James J. Little · 2017
Cited alongside, same era.
Meaningful head movements driven by emotional synthetic speech
Najmeh Sadoughi, Yang Liu, and Carlos Busso · 2017
Cited alongside, same era.
Synthesizing obama: Learning lip sync from audio
Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman · 2017
Cited alongside, same era.
Analyzing input and output representations for speech-driven gesture generation
Taras Kucherenko, Dai Hasegawa, Gustav Eje Henter, Naoshi Kaneko, and Hedvig Kjellström · 2019
Later among the works it cites.
Talking with hands 16.2m: A large-scale dataset of synchronized body-finger motion and audio for conversational motion analysis and synthesis
G. Lee, Z. Deng, S. Ma, T. Shiratori, S.S. Srinivasa, and Y. Sheikh · 2019
Later among the works it cites.
3d human pose estimation in video with temporal convolutions and semi-supervised training
Dario Pavllo, Christoph Feichtenhofer, David Grangier, and Michael Auli · 2019
Later among the works it cites.
ffmpeg-normalize. github.com/slhck/ffmpeg-normalize, 2019
Werner Robitza · 2019
Later among the works it cites.
Speech-driven animation with meaningful behaviors
Najmeh Sadoughi and Carlos Busso · 2019
Later among the works it cites.
Synthesising 3d facial motion from ”in-the-wild” speech
Panagiotis Tzirakis, Athanasios Papaioannou, Alexander Lattas, Michail Tarasiou, Björn W. Schuller, and Stefanos Zafeiriou · 2019
Later among the works it cites.
Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots
Y. Yoon, W. Ko, M. Jang, J. Lee, J. Kim, and G. Lee · 2019
Later among the works it cites.
Style transfer for co-speech gesture animation: A multi-speaker conditional-mixture approach
Chaitanya Ahuja, Dong Won Lee, Yukiko I. Nakano, and Louis-Philippe Morency · 2020
Later among the works it cites.
Style-controllable speech-driven gesture synthesis using normalising flows
Simon Alexanderson, Gustav Eje Henter, Taras Kucherenko, and Jonas Beskow · 2020
Later among the works it cites.
Gesticulator: A framework for semantically-aware speech-driven gesture generation
Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav Eje Henter, Simon Alexandersson, Iolanda Leite, and Hedvig Kjellström · 2020
Later among the works it cites.
XNect: Real-time multi-person 3D motion capture with a single RGB camera
Dushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu, Mohamed Elgharib, Pascal Fua, Hans-Peter Seidel, Helge Rhodin, Gerard Pons-Moll, and Christian Theobalt · 2020
Later among the works it cites.
Speech gesture generation from the trimodal context of text, audio, and speaker identity
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee · 2020
Later among the works it cites.
Monocular real-time hand shape and motion capture using multi-modal data
Yuxiao Zhou, Marc Habermann, Weipeng Xu, Ikhsanul Habibie, Christian Theobalt, and Feng Xu · 2020
Later among the works it cites.