Fetching the paper…
Reading the bibliography…
People naturally conduct spontaneous body motions to enhance their speeches while giving talks.
Hand movements
Paul Ekman and Wallace V Friesen · 1972
Earlier work this paper cites.
Hand and mind: What gestures reveal about thought
Michael Studdert-Kennedy · 1994
Earlier work this paper cites.
Beat: the behavior expression animation toolkit
Justine Cassell, Hannes Högni Vilhjálmsson, and Timothy Bickmore · 2004
Earlier work this paper cites.
Gesture: Visible action as utterance
Adam Kendon · 2004
Earlier work this paper cites.
Synthesizing multimodal utterances for conversational agents
Stefan Kopp and Ipke Wachsmuth · 2004
Earlier work this paper cites.
Greta. a believable embodied conversational agent
Isabella Poggi, Catherine Pelachaud, Fiorella de Rosis, Valeria Carofiglio, and Berardina De Carolis · 2005
Earlier work this paper cites.
Gesture and thought
David McNeill · 2008
Earlier work this paper cites.
Gesture controllers
Sergey Levine, Philipp Krähenbühl, Sebastian Thrun, and Vladlen Koltun · 2010
Earlier work this paper cites.
Design, analysis and experimental evaluation of block based transformation in mfcc computation for speaker recognition
Md Sahidullah and Goutam Saha · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Gesture and speech in interaction: An overview, 2014
Petra Wagner, Zofia Malisz, and Stefan Kopp · 2014
Earlier work this paper cites.
RMPE: Regional multi-person pose estimation
Hao-Shu Fang, Shuqin Xie, Yu-Wing Tai, and Cewu Lu · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman · 2017
Cited alongside, same era.
Crowdpose: Efficient crowded scenes pose estimation and a new benchmark
Jiefeng Li, Can Wang, Hao Zhu, Yihuan Mao, Hao-Shu Fang, and Cewu Lu · 2018
Cited alongside, same era.
Audio to body dynamics
Eli Shlizerman, Lucio Dery, Hayden Schoen, and Ira Kemelmacher-Shlizerman · 2018
Cited alongside, same era.
Pose Flow: Efficient online pose tracking
Yuliang Xiu, Jiefeng Li, Haoyu Wang, Yinghong Fang, and Cewu Lu · 2018
Cited alongside, same era.
Mt-vae: Learning motion transformations to generate multimodal human dynamics
Xinchen Yan, Akash Rastogi, Ruben Villegas, Kalyan Sunkavalli, Eli Shechtman, Sunil Hadap, Ersin Yumer, and Honglak Lee · 2018
Cited alongside, same era.
Everybody dance now
Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A Efros · 2019
Talking-head generation with rhythmic head motion
Lele Chen, Guofeng Cui, Celong Liu, Zhong Li, Ziyi Kou, Yi Xu, and Chenliang Xu · 2020
Later among the works it cites.
Face x-ray for more general face forgery detection
Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo · 2020
Later among the works it cites.
Speech2video synthesis with 3d skeleton regularization and expressive body poses
Miao Liao, Sibo Zhang, Peng Wang, Hao Zhu, Xinxin Zuo, and Ruigang Yang · 2020
Later among the works it cites.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Later among the works it cites.
Audio-driven talking face video generation with learning-based personalized head pose
Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao, and Yong-Jin Liu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Text-based editing of talking-head video
Ohad Fried, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shechtman, Dan B Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala · 2019
Cited alongside, same era.
Learning individual styles of conversational gesture
Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik · 2019
Cited alongside, same era.
Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots
Youngwoo Yoon, Woo-Ri Ko, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee · 2019
Cited alongside, same era.
Style transfer for co-speech gesture animation: A multi-speaker conditional-mixture approach
Chaitanya Ahuja, Dong Won Lee, Yukiko I Nakano, and Louis-Philippe Morency · 2020
Cited alongside, same era.
A stochastic conditioning scheme for diverse human motion prediction
Sadegh Aliakbarian, Fatemeh Sadat Saleh, Mathieu Salzmann, Lars Petersson, and Stephen Gould · 2020
Cited alongside, same era.
What comprises a good talking-head video generation?: A survey and benchmark
Lele Chen, Guofeng Cui, Ziyi Kou, Haitian Zheng, and Chenliang Xu · 2020
Cited alongside, same era.
Speech gesture generation from the trimodal context of text, audio, and speaker identity
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee · 2020
Later among the works it cites.
Makelttalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li · 2020
Later among the works it cites.
Magdr: Mask-guided detection and reconstruction for defending deepfakes
Zhikai Chen, Lingxi Xie, Shanmin Pang, Yong He, and Bo Zhang · 2021
Later among the works it cites.
Ai choreographer: Music conditioned 3d dance generation with aist++
Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa · 2021
Later among the works it cites.
Speech drives templates: Co-speech gesture synthesis with learned templates
Shenhan Qian, Zhi Tu, Yihao Zhi, Wen Liu, and Shenghua Gao · 2021
Later among the works it cites.
Multi-attentional deepfake detection
Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei, Weiming Zhang, and Nenghai Yu · 2021
Later among the works it cites.