Fetching the paper…
Reading the bibliography…
This work addresses the problem of generating 3D holistic body motions from human speech.
The role of gesture in communication and thinking
Susan Goldin-Meadow · 1999
Earlier work this paper cites.
Beat: the behavior expression animation toolkit
Justine Cassell, Hannes Högni Vilhjálmsson, and Timothy Bickmore · 2001
Earlier work this paper cites.
Gesture: Visible action as utterance
Adam Kendon · 2004
Earlier work this paper cites.
Synthesizing multimodal utterances for conversational agents
Stefan Kopp and Ipke Wachsmuth · 2004
Earlier work this paper cites.
Greta. a believable embodied conversational agent
Isabella Poggi, Catherine Pelachaud, F de Rosis, Valeria Carofiglio, and B De Carolis · 2005
Earlier work this paper cites.
Nonlinear equations
Jorge Nocedal and Stephen J Wright · 2006
Earlier work this paper cites.
Real-time prosody-driven synthesis of body language
Sergey Levine, Christian Theobalt, and Vladlen Koltun · 2009
Earlier work this paper cites.
A 3-D Audio-Visual Corpus of Affective Communication
Gabriele Fanelli, Juergen Gall, Harald Romsdorfer, Thibaut Weise, and Luc Van Gool · 2010
Earlier work this paper cites.
Gesture controllers
Sergey Levine, Philipp Krähenbühl, Sebastian Thrun, and Vladlen Koltun · 2010
Earlier work this paper cites.
Design, analysis and experimental evaluation of block based transformation in mfcc computation for speaker recognition
Md Sahidullah and Goutam Saha · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, Andrew Y Ng, et al · 2013
Earlier work this paper cites.
Virtual character performance from speech
Stacy Marsella, Yuyu Xu, Margaux Lhommet, Andrew Feng, Stefan Scherer, and Ari Shapiro · 2013
Earlier work this paper cites.
Gesture and speech in interaction: An overview, 2014
Petra Wagner, Zofia Malisz, and Stefan Kopp · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Keep it smpl: Automatic estimation of 3d human pose and shape from a single image
Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J Black · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Earlier work this paper cites.
Realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2017
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation
Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam · 2017
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Earlier work this paper cites.
Meaningful head movements driven by emotional synthetic speech
Najmeh Sadoughi, Yang Liu, and Carlos Busso · 2017
Earlier work this paper cites.
Creating a gesture-speech dataset for speech-based automatic gesture generation
Kenta Takeuchi, Souichirou Kubota, Keisuke Suzuki, Dai Hasegawa, and Hiroshi Sakuta · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Investigating the use of recurrent motion modelling for speech gesture generation
Ylva Ferstl and Rachel McDonnell · 2018
Cited alongside, same era.
Total capture: A 3d deformation model for tracking faces, hands, and bodies
Hanbyul Joo, Tomas Simon, and Yaser Sheikh · 2018
Cited alongside, same era.
End-to-end recovery of human shape and pose
Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik · 2018
Cited alongside, same era.
Visemenet: Audio-driven animator-centric speech animation
Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh · 2018
Cited alongside, same era.
Capture, learning, and synthesis of 3D speaking styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael Black · 2019
Speech2affectivegestures: Synthesizing co-speech gestures with generative adversarial affective expression learning
Uttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, and Dinesh Manocha · 2021
Later among the works it cites.
Exploring simple siamese representation learning
Xinlei Chen and Kaiming He · 2021
Later among the works it cites.
Collaborative regression of expressive bodies using moderation
Yao Feng, Vasileios Choutas, Timo Bolkart, Dimitrios Tzionas, and Michael J Black · 2021
Later among the works it cites.
Learning speech-driven 3d conversational gestures from video
Ikhsanul Habibie, Weipeng Xu, Dushyant Mehta, Lingjie Liu, Hans-Peter Seidel, Gerard Pons-Moll, Mohamed Elgharib, and Christian Theobalt · 2021
Later among the works it cites.
Audio2gestures: Generating diverse gestures from speech audio with conditional variational autoencoders
Jing Li, Di Kang, Wenjie Pei, Xuefei Zhe, Ying Zhang, Zhenyu He, and Linchao Bao · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning individual styles of conversational gesture
S. Ginosar, A. Bar, G. Kohavi, C. Chan, A. Owens, and J. Malik · 2019
Cited alongside, same era.
Real-time facial surface geometry from monocular video on mobile gpus
Yury Kartynnik, Artsiom Ablavatski, Ivan Grishchenko, and Matthias Grundmann · 2019
Cited alongside, same era.
Analyzing input and output representations for speech-driven gesture generation
Taras Kucherenko, Dai Hasegawa, Gustav Eje Henter, Naoshi Kaneko, and Hedvig Kjellström · 2019
Cited alongside, same era.
Expressive body capture: 3d hands, face, and body from a single image
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black · 2019
Cited alongside, same era.
Mmface: A multi-metric regression network for unconstrained face reconstruction
Hongwei Yi, Chen Li, Qiong Cao, Xiaoyong Shen, Sheng Li, Guoping Wang, and Yu-Wing Tai · 2019
Cited alongside, same era.
Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots
Youngwoo Yoon, Woo-Ri Ko, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee · 2019
Cited alongside, same era.
Smplpix: Neural avatars from 3d human models
Sergey Prokudin, Michael J Black, and Javier Romero · 2021
Later among the works it cites.
Speech drives templates: Co-speech gesture synthesis with learned templates
Shenhan Qian, Zhi Tu, Yihao Zhi, Wen Liu, and Shenghua Gao · 2021
Later among the works it cites.
Meshtalk: 3d face animation from speech using cross-modality disentanglement
Alexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre, and Yaser Sheikh · 2021
Later among the works it cites.
Pymaf: 3d human pose and shape regression with pyramidal mesh alignment feedback loop
Hongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang, Yebin Liu, Limin Wang, and Zhenan Sun · 2021
Later among the works it cites.
Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings
Tenglong Ao, Qingzhe Gao, Yuke Lou, Baoquan Chen, and Libin Liu · 2022
Closest in time.
Faceformer: Speech-driven 3d facial animation with transformers
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura · 2022
Closest in time.
Learning hierarchical cross-modal association for co-speech gesture generation
Xian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu, Rui Qian, Xinyi Lin, Xiaowei Zhou, Wayne Wu, Bo Dai, and Bolei Zhou · 2022
Closest in time.
Learning to listen: Modeling non-deterministic dyadic facial motion
Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li, Trevor Darrell, Angjoo Kanazawa, and Shiry Ginosar · 2022
Closest in time.
Bailando: 3d dance generation by actor-critic gpt with choreographic memory
Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu · 2022
Closest in time.
Multiface: A dataset for neural face rendering
Cheng-hsin Wuu, Ningyuan Zheng, Scott Ardisson, Rohan Bali, Danielle Belko, Eric Brockmeyer, Lucas Evans, Timothy Godisart, Hyowon Ha, Alexander Hypes, Taylor Koska, Steven Krenn, Stephen Lombardi, Xiaomin Luo, Kevyn McPhail, Laura Millerschoen, Michal Perdoch, Mark Pitts, Alexander Richard, Jason Saragih, Junko Saragih, Takaaki Shiratori, Tomas Simon, Matt Stewart, Autumn Trimble, Xinshuo Weng, David Whitewolf, Chenglei Wu, Shoou-I Yu, and Yaser Sheikh · 2022
Closest in time.
Freeform body motion generation from speech
Jing Xu, Wei Zhang, Yalong Bai, Qibin Sun, and Tao Mei · 2022
Closest in time.
Gesture2vec: Clustering gestures using representation learning methods for co-speech gesture generation
Payam Jome Yazdian, Mo Chen, and Angelica Lim · 2022
Closest in time.
Human-aware object placement for visual environment reconstruction
Hongwei Yi, Chun-Hao P. Huang, Dimitrios Tzionas, Muhammed Kocabas, Mohamed Hassan, Siyu Tang, Justus Thies, and Michael J. Black · 2022
Closest in time.
Pymaf-x: Towards well-aligned full-body model regression from monocular images
Hongwen Zhang, Yating Tian, Yuxiang Zhang, Mengcheng Li, Liang An, Zhenan Sun, and Yebin Liu · 2022
Closest in time.
Towards metrical reconstruction of human faces
Wojciech Zielonka, Timo Bolkart, and Justus Thies · 2022
Closest in time.
Learning analytical posterior probability for human mesh recovery
Qi Fang, Kang Chen, Yinghui Fan, Qing Shuai, Jiefeng Li, and Weidong Zhang · 2023
Closest in time.
NIKI: Neural inverse kinematics with invertible neural networks for 3d human pose and shape estimation
Jiefeng Li, Siyuan Bian, Qi Liu, Jiasheng Tang, Fan Wang, and Cewu Lu · 2023
Closest in time.
Trace: Temporal regression of 5d avatars with dynamic cameras in 3d environments
Yu Sun, Qian Bao, Wu Liu, Tao Mei, and Michael J Black · 2023
Closest in time.