Fetching the paper…
Reading the bibliography…
Automatic synthesis of realistic co-speech gestures is an increasingly important yet challenging task in artificial embodied agent creation.
The repertoire of nonverbal behavior: Categories, origins, usage, and coding
Paul Ekman and Wallace V Friesen. 1969 · 1969
Earlier work this paper cites.
Nonverbal behaviors, persuasion, and credibility
Judee K Burgoon, Thomas Birk, and Michael Pfau. 1990 · 1990
Earlier work this paper cites.
Hand and Mind
David McNeill. 1992 · 1992
Earlier work this paper cites.
Gestural beats: The rhythm hypothesis
Evelyn McClave. 1994 · 1994
Earlier work this paper cites.
Linguistic Features of Metaphoric Gestures
Rebecca A. Webb. 1996 · 1996
Earlier work this paper cites.
Motion Graphs
Lucas Kovar, Michael Gleicher, and Frédéric Pighin. 2002 · 2002
Earlier work this paper cites.
Beat: the behavior expression animation toolkit
Justine Cassell, Hannes Högni Vilhjálmsson, and Timothy Bickmore. 2004 · 2004
Earlier work this paper cites.
Gesture: Visible Action as Utterance
Adam Kendon. 2004 · 2004
Earlier work this paper cites.
Gesture Generation by Imitation: From Human Behavior to Computer Character Animation
Michael Kipp. 2004 · 2004
Earlier work this paper cites.
A Tutorial on Onset Detection in Music Signals
J.P. Bello, L. Daudet, S. Abdallah, C. Duxbury, M. Davies, and M.B. Sandler. 2005 · 2005
Earlier work this paper cites.
Making Them Dance.. In AAAI Fall Symposium: Aurally Informed Performance , Vol. 2
Jae Woo Kim, Hesham Fouad, and James K Hahn. 2006 · 2006
Earlier work this paper cites.
Towards a common framework for multimodal generation: The behavior markup language. In International workshop on intelligent virtual agents . Springer, 205–217
Stefan Kopp, Brigitte Krenn, Stacy Marsella, Andrew N Marshall, Catherine Pelachaud, Hannes Pirker, Kristinn R Thórisson, and Hannes Vilhjálmsson. 2006 · 2006
Earlier work this paper cites.
Beat tracking by dynamic programming
Daniel PW Ellis. 2007 · 2007
Earlier work this paper cites.
Gesture Modeling and Animation Based on a Probabilistic Re-Creation of Speaker Style
Michael Neff, Michael Kipp, Irene Albrecht, and Hans-Peter Seidel. 2008 · 2008
Earlier work this paper cites.
Gesture Controllers
Sergey Levine, Philipp Krähenbühl, Sebastian Thrun, and Vladlen Koltun. 2010 · 2010
Earlier work this paper cites.
Robot behavior toolkit: generating effective social behaviors for robots. In 2012 7th ACM/IEEE International Conference on Human-Robot Interaction (HRI) . IEEE, 25–32
Chien-Ming Huang and Bilge Mutlu. 2012 · 2012
Earlier work this paper cites.
Temporal, Structural, and Pragmatic Synchrony between Intonation and Gesture
Daniel P. Loehr. 2012 · 2012
Earlier work this paper cites.
Extraction and alignment evaluation of motion beats for street dance. In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2429–2433
Chieh Ho, Wei-Tze Tsai, Keng-Sheng Lin, and Homer H Chen. 2013 · 2013
Earlier work this paper cites.
Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. 2013 · 2013
Earlier work this paper cites.
Gesture and speech in interaction: An overview
Petra Wagner, Zofia Malisz, and Stefan Kopp. 2014 · 2014
Earlier work this paper cites.
Predicting co-verbal gestures: A deep and temporal modeling approach. In International Conference on Intelligent Virtual Agents . Springer, 152–166
Chung-Cheng Chiu, Louis-Philippe Morency, and Stacy Marsella. 2015 · 2015
Earlier work this paper cites.
Librosa: Audio and Music Signal Analysis in Python
Brian McFee, Colin Raffel, Dawen Liang, Daniel P. W. Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. 2015 · 2015
Earlier work this paper cites.
Naoqi api documentation. In 2016 IEEE International Conference on Multimedia and Expo (ICME), vol. http://doc. aldebaran. com/2-5/homepepper. html
Robotics Softbank. 2018 · 2016
Earlier work this paper cites.
Convolutional Pose Machines. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Shih-En Wei, Varun Ramakrishna, Takeo Kanade, and Yaser Sheikh. 2016 · 2016
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
Predicting head pose from speech with a conditional variational autoencoder. ISCA
David Greenwood, Stephen Laycock, and Iain Matthews. 2017 · 2017
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Cited alongside, same era.
Phase-functioned neural networks for character control
Daniel Holden, Taku Komura, and Jun Saito. 2017 · 2017
Cited alongside, same era.
Categorical Reparameterization with Gumbel-Softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2017 · 2017
Cited alongside, same era.
Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi.. In Interspeech , Vol. 2017. 498–502
Jukebox: A Generative Model for Music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. 2020 · 2020
Later among the works it cites.
Robust Motion In-Betweening
Félix G. Harvey, Mike Yurick, Derek Nowrouzezahrai, and Christopher Pal. 2020 · 2020
Later among the works it cites.
Moglow: Probabilistic and controllable motion synthesis using normalising flows
Gustav Eje Henter, Simon Alexanderson, and Jonas Beskow. 2020 · 2020
Later among the works it cites.
Gesticulator: A framework for semantically-aware speech-driven gesture generation. In Proceedings of the 2020 International Conference on Multimodal Interaction . 242–250
Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav Eje Henter, Simon Alexandersson, Iolanda Leite, and Hedvig Kjellström. 2020 · 2020
Later among the works it cites.
A Large, Crowdsourced Evaluation of Gesture Generation Systems on Common Data: The GENEA Challenge 2020. In 26th International Conference on Intelligent User Interfaces (College Station, TX, USA) (IUI ’21) . Association for Computing Machinery, New York, NY, USA, 11–21
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. 2017 · 2017
Cited alongside, same era.
Vnect: Real-time 3d human pose estimation with a single rgb camera
Dushyant Mehta, Srinath Sridhar, Oleksandr Sotnychenko, Helge Rhodin, Mohammad Shafiei, Hans-Peter Seidel, Weipeng Xu, Dan Casas, and Christian Theobalt. 2017 · 2017
Cited alongside, same era.
Neural Discrete Representation Learning
van den Aaron Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Cited alongside, same era.
Attention Is All You Need. In Advances in Neural Information Processing Systems , Vol. 30. Curran Associates, Inc
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Deep motifs and motion signatures
Andreas Aristidou, Daniel Cohen-Or, Jessica K Hodgins, Yiorgos Chrysanthou, and Ariel Shamir. 2018 · 2018
Cited alongside, same era.
Visual rhythm and beat. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops . 2532–2535
Abe Davis and Maneesh Agrawala. 2018 · 2018
Cited alongside, same era.
Investigating the use of recurrent motion modelling for speech gesture generation. In Proceedings of the 18th International Conference on Intelligent Virtual Agents . 93–98
Ylva Ferstl and Rachel McDonnell. 2018 · 2018
Cited alongside, same era.
Taras Kucherenko, Patrik Jonell, Youngwoo Yoon, Pieter Wolfert, and Gustav Eje Henter. 2021a · 2020
Later among the works it cites.
Character Controllers Using Motion VAEs
Hung Yu Ling, Fabio Zinno, George Cheng, and Michiel van de Panne. 2020 · 2020
Later among the works it cites.
Local motion phases for learning multi-contact character movements
Sebastian Starke, Yiwei Zhao, Taku Komura, and Kazi Zaman. 2020 · 2020
Later among the works it cites.
Speech gesture generation from the trimodal context of text, audio, and speaker identity
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee. 2020 · 2020
Later among the works it cites.
Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents
Uttaran Bhattacharya, Nicholas Rewkowski, Abhishek Banerjee, Pooja Guhan, Aniket Bera, and Dinesh Manocha. 2021b · 2021
Later among the works it cites.
Choreomaster: choreography-oriented music-driven dance synthesis
Kang Chen, Zhipeng Tan, Jin Lei, Song-Hai Zhang, Yuan-Chen Guo, Weidong Zhang, and Shi-Min Hu. 2021 · 2021
Later among the works it cites.
Learning speech-driven 3d conversational gestures from video. In Proceedings of the 21st ACM International Conference on Intelligent Virtual Agents . 101–108
Ikhsanul Habibie, Weipeng Xu, Dushyant Mehta, Lingjie Liu, Hans-Peter Seidel, Gerard Pons-Moll, Mohamed Elgharib, and Christian Theobalt. 2021 · 2021
Later among the works it cites.
Speech Drives Templates: Co-Speech Gesture Synthesis with Learned Templates. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 11077–11086
Shenhan Qian, Zhi Tu, Yihao Zhi, Wen Liu, and Shenghua Gao. 2021 · 2021
Later among the works it cites.
Zero-Shot Text-to-Image Generation. In Proceedings of the 38th International Conference on Machine Learning . PMLR, 8821–8831
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement
Alexander Richard, Michael Zollhoefer, Yandong Wen, de la Fernando Torre, and Yaser Sheikh. 2021 · 2021
Later among the works it cites.
Transflower: probabilistic autoregressive dance generation with multimodal attention
Guillermo Valle-Pérez, Gustav Eje Henter, Jonas Beskow, Andre Holzapfel, Pierre-Yves Oudeyer, and Simon Alexanderson. 2021 · 2021
Later among the works it cites.
Videogpt: Video generation using vq-vae and transformers
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. 2021 · 2021
Later among the works it cites.
Rhythm is a Dancer: Music-Driven Motion Synthesis with Global Structure
Andreas Aristidou, Anastasios Yiannakidis, Kfir Aberman, Daniel Cohen-Or, Ariel Shamir, and Yiorgos Chrysanthou. 2022 · 2022
Closest in time.
Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Xian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu, Rui Qian, Xinyi Lin, Xiaowei Zhou, Wayne Wu, Bo Dai, and Bolei Zhou. 2022 · 2022
Closest in time.
Bailando: 3D Dance Generation by Actor-Critic GPT With Choreographic Memory. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 11050–11059
Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. 2022 · 2022
Closest in time.
DeepPhase: Periodic Autoencoders for Learning Motion Phase Manifolds
Sebastian Starke, Ian Mason, and Taku Komura. 2022 · 2022
Closest in time.
A Review of Evaluation Practices of Gesture Generation in Embodied Conversational Agents
Pieter Wolfert, Nicole Robinson, and Tony Belpaeme. 2022 · 2022
Closest in time.
Freeform Body Motion Generation from Speech
Jing Xu, Wei Zhang, Yalong Bai, Qibin Sun, and Tao Mei. 2022 · 2022
Closest in time.
Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning. In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21) . Association for Computing Machinery, New York, NY, USA, 2027–2036
Uttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, and Dinesh Manocha. 2021a · 2036
Closest in time.