Fetching the paper…
Reading the bibliography…
Generating speech-consistent body and gesture movements is a long-standing problem in virtual avatar creation.
Robust estimation of a location parameter
Peter J Huber · 1992
Earlier work this paper cites.
Animated conversation: rule-based generation of facial expression, gesture & spoken intonation for multiple conversational agents
Justine Cassell, Catherine Pelachaud, Norman Badler, Mark Steedman, Brett Achorn, Tripp Becket, Brett Douville, Scott Prevost, and Matthew Stone · 1994
Earlier work this paper cites.
The persona effect: how substantial is it?
Susanne Van Mulken, Elisabeth Andre, and Jochen Müller · 1998
Earlier work this paper cites.
Speech-gesture mismatches: Evidence for one underlying representation of linguistic and nonlinguistic information
Justine Cassell, David McNeill, and Karl-Erik McCullough · 1999
Earlier work this paper cites.
Learning statistical models of human motion
Richard Bowden · 2000
Earlier work this paper cites.
Style machines
Matthew Brand and Aaron Hertzmann · 2000
Earlier work this paper cites.
Animating by multi-level sampling
Katherine Pullen and Christoph Bregler · 2000
Earlier work this paper cites.
Learning variable-length markov models of behavior
Aphrodite Galata, Neil Johnson, and David Hogg · 2001
Earlier work this paper cites.
Dynamic bayesian networks for audio-visual speech recognition
Ara V Nefian, Luhong Liang, Xiaobo Pi, Xiaoxing Liu, and Kevin Murphy · 2002
Earlier work this paper cites.
Beat: the behavior expression animation toolkit
Justine Cassell, Hannes Högni Vilhjálmsson, and Timothy Bickmore · 2004
Earlier work this paper cites.
Beat tracking by dynamic programming
Daniel Ellis · 2007
Earlier work this paper cites.
Gesture controllers
Sergey Levine, Philipp Krähenbühl, Sebastian Thrun, and Vladlen Koltun · 2010
Earlier work this paper cites.
Hand and mind
David McNeill · 2011
Earlier work this paper cites.
A friendly gesture: Investigating the effect of multimodal robot behavior in human-robot interaction
Maha Salem, Katharina Rohlfing, Stefan Kopp, and Frank Joublin · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
Robot behavior toolkit: generating effective social behaviors for robots
Chien-Ming Huang and Bilge Mutlu · 2012
Earlier work this paper cites.
Temporal, structural, and pragmatic synchrony between intonation and gesture
Daniel P Loehr · 2012
Earlier work this paper cites.
Generation and evaluation of communicative robot gesture
Maha Salem, Stefan Kopp, Ipke Wachsmuth, Katharina Rohlfing, and Frank Joublin · 2012
Earlier work this paper cites.
Virtual character performance from speech
Stacy Marsella, Yuyu Xu, Margaux Lhommet, Andrew Feng, Stefan Scherer, and Ari Shapiro · 2013
Earlier work this paper cites.
Gesture and speech in interaction: An overview
P. Wagner, Z. Malisz, and S. Kopp · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Msp-avatar corpus: Motion capture recordings to study the role of discourse functions in the design of intelligent virtual agents
Najmeh Sadoughi, Yang Liu, and Carlos Busso · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
A deep learning framework for character motion synthesis and editing
Daniel Holden, Jun Saito, and Taku Komura · 2016
Earlier work this paper cites.
Stacked hourglass networks for human pose estimation
Alejandro Newell, Kaiyu Yang, and Jia Deng · 2016
Earlier work this paper cites.
Gentle: A forced aligner
Ochshorn Robert and Hawkin Max · 2016
Earlier work this paper cites.
Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network
Wenzhe Shi, Jose Caballero, Ferenc Huszar, Johannes Totz, Andrew P. Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang · 2016
Earlier work this paper cites.
A multimodal motion-captured corpus of matched and mismatched extravert-introvert conversational pairs
Jackson Tolins, Kris Liu, Yingying Wang, Jean E Fox Tree, Marilyn Walker, and Michael Neff · 2016
Earlier work this paper cites.
Learning human motion models for long-term predictions
Partha Ghosh, Jie Song, Emre Aksan, and Otmar Hilliges · 2017
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Cited alongside, same era.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2018
Cited alongside, same era.
Hp-gan: Probabilistic 3d human motion prediction via gan
Emad Barsoum, John Kender, and Zicheng Liu · 2018
Cited alongside, same era.
Monocular expressive body regression through body-driven attention
Vasileios Choutas, Georgios Pavlakos, Timo Bolkart, Dimitrios Tzionas, and Michael J. Black · 2020
Later among the works it cites.
In defence of metric learning for speaker recognition
Joon Son Chung, Jaesung Huh, Seongkyu Mun, Minjae Lee, Hee Soo Heo, Soyeon Choe, Chiheon Ham, Sunghwan Jung, Bong-Jin Lee, and Icksang Han · 2020
Later among the works it cites.
Adversarial gesture generation with realistic gesture phasing
Ylva Ferstl, Michael Neff, and Rachel McDonnell · 2020
Later among the works it cites.
Dance revolution: Long-term dance generation with music via curriculum learning
Ruozi Huang, Huang Hu, Wei Wu, Kei Sawada, Mi Zhang, and Daxin Jiang · 2020
Later among the works it cites.
Gesticulator: A framework for semantically-aware speech-driven gesture generation
Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav Eje Henter, Simon Alexandersson, Iolanda Leite, and Hedvig Kjellström · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dai Hasegawa, Naoshi Kaneko, Shinichi Shirakawa, Hiroshi Sakuta, and Kazuhiko Sumi · 2018
Cited alongside, same era.
A speech-driven hand gesture generation method and evaluation in android robots
Carlos T. Ishi, Daichi Machiyashiki, Ryusuke Mikata, and Hiroshi Ishiguro · 2018
Cited alongside, same era.
Convolutional sequence to sequence model for human dynamics
C. Li, Z. Zhang, W. S. Lee, and G. H. Lee · 2018
Cited alongside, same era.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Cited alongside, same era.
Neural kinematic networks for unsupervised motion retargetting
Ruben Villegas, Jimei Yang, Duygu Ceylan, and Honglak Lee · 2018
Cited alongside, same era.
Language2pose: Natural language grounded pose forecasting
Chaitanya Ahuja and Louis-Philippe Morency · 2019
Cited alongside, same era.
Structured prediction helps 3d human motion modelling
Emre Aksan, Manuel Kaufmann, and Otmar Hilliges · 2019
Cited alongside, same era.
Later among the works it cites.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Later among the works it cites.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Yapeng Tian, Dingzeyu Li, and Chenliang Xu · 2020
Later among the works it cites.
Deep high-resolution representation learning for visual recognition
Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, et al · 2020
Later among the works it cites.
Lightweight and efficient end-to-end speech recognition using low-rank transformer
Genta Indra Winata, Samuel Cahyawijaya, Zhaojiang Lin, Zihan Liu, and Pascale Fung · 2020
Later among the works it cites.
Speech gesture generation from the trimodal context of text, audio, and speaker identity
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee · 2020
Later among the works it cites.
Audio-visual recognition of overlapped speech for the lrs2 dataset
Jianwei Yu, Shi-Xiong Zhang, Jian Wu, Shahram Ghorbani, Bo Wu, Shiyin Kang, Shansong Liu, Xunying Liu, Helen Meng, and Dong Yu · 2020
Later among the works it cites.
Sep-stereo: Visually guided stereophonic audio generation by associating source separation
Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu · 2020
Later among the works it cites.
Makelttalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li · 2020
Later among the works it cites.
Generative tweening: Long-term inbetweening of 3d human motions
Yi Zhou, Jingwan Lu, Connelly Barnes, Jimei Yang, Sitao Xiang, et al · 2020
Later among the works it cites.
Glocalnet: Class-aware long-term human motion synthesis
Neeraj Battan, Yudhik Agrawal, Sai Soorya Rao, Aman Goel, and Avinash Sharma · 2021
Later among the works it cites.
Text2gestures: A transformer-based network for generating emotive body gestures for virtual agents
Uttaran Bhattacharya, Nicholas Rewkowski, Abhishek Banerjee, Pooja Guhan, Aniket Bera, and Dinesh Manocha · 2021
Later among the works it cites.
Learning speech-driven 3d conversational gestures from video
Ikhsanul Habibie, Weipeng Xu, Dushyant Mehta, Lingjie Liu, Hans-Peter Seidel, Gerard Pons-Moll, Mohamed Elgharib, and Christian Theobalt · 2021
Later among the works it cites.
Audio-driven emotional video portraits
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu · 2021
Later among the works it cites.
Dancenet3d: Music based dance generation with parametric motion transformer
Buyu Li, Yongchi Zhao, and Lu Sheng · 2021
Later among the works it cites.
Jing Li, Di Kang, Wenjie Pei, Xuefei Zhe, Ying Zhang, Zhenyu He, and Linchao Bao · 2021
Later among the works it cites.
Learn to dance with aist++: Music conditioned 3d dance generation
Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa · 2021
Later among the works it cites.
Speech drives templates: Co-speech gesture synthesis with learned templates
Shenhan Qian, Zhi Tu, YiHao Zhi, Wen Liu, and Shenghua Gao · 2021
Later among the works it cites.
Cyclic co-learning of sounding object visual grounding and sound separation
Yapeng Tian, Di Hu, and Chenliang Xu · 2021
Later among the works it cites.
Visually informed binaural audio generation without binaural audios
Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, and Dahua Lin · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Later among the works it cites.
Visual sound localization in the wild by cross-modal interference erasing
Xian Liu, Rui Qian, Hang Zhou, Di Hu, Weiyao Lin, Ziwei Liu, Bolei Zhou, and Xiaowei Zhou · 2022
Closest in time.
Semantic-aware implicit neural audio-driven video portrait generation
Xian Liu, Yinghao Xu, Qianyi Wu, Hang Zhou, Wayne Wu, and Bolei Zhou · 2022
Closest in time.