Fetching the paper…
Reading the bibliography…
Co-speech gesture is crucial for human-machine interaction and digital entertainment.
Nonverbal behaviors, persuasion, and credibility
Judee K Burgoon, Thomas Birk, and Michael Pfau · 1990
Earlier work this paper cites.
Animated conversation: rule-based generation of facial expression, gesture & spoken intonation for multiple conversational agents
Justine Cassell, Catherine Pelachaud, Norman Badler, Mark Steedman, Brett Achorn, Tripp Becket, Brett Douville, Scott Prevost, and Matthew Stone · 1994
Earlier work this paper cites.
Hand and mind: What gestures reveal about thought
Michael Studdert-Kennedy · 1994
Earlier work this paper cites.
Numerical linear algebra
Lloyd N Trefethen and David Bau III · 1997
Earlier work this paper cites.
The persona effect: how substantial is it?
Susanne Van Mulken, Elisabeth Andre, and Jochen Müller · 1998
Earlier work this paper cites.
Speech-gesture mismatches: Evidence for one underlying representation of linguistic and nonlinguistic information
Justine Cassell, David McNeill, and Karl-Erik McCullough · 1999
Earlier work this paper cites.
Beat: the behavior expression animation toolkit
Justine Cassell, Hannes Högni Vilhjálmsson, and Timothy Bickmore · 2004
Earlier work this paper cites.
Gesture: Visible action as utterance
Adam Kendon · 2004
Earlier work this paper cites.
Hand and mind
David McNeill · 2011
Earlier work this paper cites.
The interplay between gesture and speech in the production of referring expressions: Investigating the tradeoff hypothesis
Jan P De Ruiter, Adrian Bangerter, and Paula Dings · 2012
Earlier work this paper cites.
Displaced dynamic expression regression for real-time facial tracking and animation
Chen Cao, Qiming Hou, and Kun Zhou · 2014
Earlier work this paper cites.
The mpi emotional body expressions database for narrative scenarios
Ekaterina Volkova, Stephan De La Rosa, Heinrich H Bülthoff, and Betty Mohler · 2014
Earlier work this paper cites.
Gesture and speech in interaction: An overview
P. Wagner, Z. Malisz, and S. Kopp · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
librosa: Audio and music signal analysis in python, 2015
McFee, Brian, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto · 2015
Earlier work this paper cites.
Msp-avatar corpus: Motion capture recordings to study the role of discourse functions in the design of intelligent virtual agents
Najmeh Sadoughi, Yang Liu, and Carlos Busso · 2015
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei · 2016
Earlier work this paper cites.
Face2face: Real-time face capture and reenactment of rgb videos
Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner · 2016
Earlier work this paper cites.
A multimodal motion-captured corpus of matched and mismatched extravert-introvert conversational pairs
Jackson Tolins, Kris Liu, Yingying Wang, Jean E Fox Tree, Marilyn Walker, and Michael Neff · 2016
Earlier work this paper cites.
In defense of the triplet loss for person re-identification
Alexander Hermans, Lucas Beyer, and Bastian Leibe · 2017
Earlier work this paper cites.
Creating a gesture-speech dataset for speech-based automatic gesture generation
Kenta Takeuchi, Souichirou Kubota, Keisuke Suzuki, Dai Hasegawa, and Hiroshi Sakuta · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Hand gestures and verbal acknowledgments improve human-robot rapport
Jason R Wilson, Nah Young Lee, Annie Saechao, Sharon Hershenson, Matthias Scheutz, and Linda Tickle-Degnen · 2017
Earlier work this paper cites.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2018
Cited alongside, same era.
Synthesizing images of humans in unseen poses
Guha Balakrishnan, Amy Zhao, Adrian V Dalca, Fredo Durand, and John Guttag · 2018
Cited alongside, same era.
Investigating the use of recurrent motion modelling for speech gesture generation
Ylva Ferstl and Rachel McDonnell · 2018
Cited alongside, same era.
Evaluation of speech-to-gesture generation using bi-directional lstm network
Dai Hasegawa, Naoshi Kaneko, Shinichi Shirakawa, Hiroshi Sakuta, and Kazuhiko Sumi · 2018
Cited alongside, same era.
A speech-driven hand gesture generation method and evaluation in android robots
Carlos T. Ishi, Daichi Machiyashiki, Ryusuke Mikata, and Hiroshi Ishiguro · 2018
Cited alongside, same era.
Monocular expressive body regression through body-driven attention
Vasileios Choutas, Georgios Pavlakos, Timo Bolkart, Dimitrios Tzionas, and Michael J. Black · 2020
Later among the works it cites.
Adversarial gesture generation with realistic gesture phasing
Ylva Ferstl, Michael Neff, and Rachel McDonnell · 2020
Later among the works it cites.
Gesticulator: A framework for semantically-aware speech-driven gesture generation
Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav Eje Henter, Simon Alexandersson, Iolanda Leite, and Hedvig Kjellström · 2020
Later among the works it cites.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Later among the works it cites.
Lightweight and efficient end-to-end speech recognition using low-rank transformer
Genta Indra Winata, Samuel Cahyawijaya, Zhaojiang Lin, Zihan Liu, and Pascale Fung · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dance with melody: An lstm-autoencoder approach to music-oriented dance synthesis
Taoran Tang, Jia Jia, and Hanyang Mao · 2018
Cited alongside, same era.
X2face: A network for controlling face generation using images, audio, and pose codes
Olivia Wiles, A Koepke, and Andrew Zisserman · 2018
Cited alongside, same era.
Language2pose: Natural language grounded pose forecasting
Chaitanya Ahuja and Louis-Philippe Morency · 2019
Cited alongside, same era.
Openpose: realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2019
Cited alongside, same era.
Everybody dance now
Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A Efros · 2019
Cited alongside, same era.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Cited alongside, same era.
Capture, learning, and synthesis of 3d speaking styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black · 2019
Cited alongside, same era.
Statistics-based motion synthesis for social conversations
Yanzhe Yang, Jimei Yang, and Jessica Hodgins · 2020
Later among the works it cites.
Speech gesture generation from the trimodal context of text, audio, and speaker identity
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee · 2020
Later among the works it cites.
Audio-visual recognition of overlapped speech for the lrs2 dataset
Jianwei Yu, Shi-Xiong Zhang, Jian Wu, Shahram Ghorbani, Bo Wu, Shiyin Kang, Shansong Liu, Xunying Liu, Helen Meng, and Dong Yu · 2020
Later among the works it cites.
Text2gestures: A transformer-based network for generating emotive body gestures for virtual agents
Uttaran Bhattacharya, Nicholas Rewkowski, Abhishek Banerjee, Pooja Guhan, Aniket Bera, and Dinesh Manocha · 2021
Later among the works it cites.
Learning speech-driven 3d conversational gestures from video
Ikhsanul Habibie, Weipeng Xu, Dushyant Mehta, Lingjie Liu, Hans-Peter Seidel, Gerard Pons-Moll, Mohamed Elgharib, and Christian Theobalt · 2021
Later among the works it cites.
Jing Li, Di Kang, Wenjie Pei, Xuefei Zhe, Ying Zhang, Zhenyu He, and Linchao Bao · 2021
Later among the works it cites.
Learn to dance with aist++: Music conditioned 3d dance generation
Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa · 2021
Later among the works it cites.
Speech drives templates: Co-speech gesture synthesis with learned templates
Shenhan Qian, Zhi Tu, YiHao Zhi, Wen Liu, and Shenghua Gao · 2021
Later among the works it cites.
Meshtalk: 3d face animation from speech using cross-modality disentanglement
Alexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre, and Yaser Sheikh · 2021
Later among the works it cites.
Motion representations for articulated animation
Aliaksandr Siarohin, Oliver J Woodford, Jian Ren, Menglei Chai, and Sergey Tulyakov · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Later among the works it cites.
Visual sound localization in the wild by cross-modal interference erasing
Xian Liu, Rui Qian, Hang Zhou, Di Hu, Weiyao Lin, Ziwei Liu, Bolei Zhou, and Xiaowei Zhou · 2022
Closest in time.
Learning hierarchical cross-modal association for co-speech gesture generation
Xian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu, Rui Qian, Xinyi Lin, Xiaowei Zhou, Wayne Wu, Bo Dai, and Bolei Zhou · 2022
Closest in time.
Semantic-aware implicit neural audio-driven video portrait generation
Xian Liu, Yinghao Xu, Qianyi Wu, Hang Zhou, Wayne Wu, and Bolei Zhou · 2022
Closest in time.
Learning to listen: Modeling non-deterministic dyadic facial motion
Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li, Trevor Darrell, Angjoo Kanazawa, and Shiry Ginosar · 2022
Closest in time.
Bailando: 3d dance generation by actor-critic gpt with choreographic memory
Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu · 2022
Closest in time.
Freeform body motion generation from speech
Jing Xu, Wei Zhang, Yalong Bai, Qibin Sun, and Tao Mei · 2022
Closest in time.
Gesture2vec: Clustering gestures using representation learning methods for co-speech gesture generation, 2022
Payam Jome Yazdian, Mo Chen, and Angelica Lim · 2022
Closest in time.