Fetching the paper…
Reading the bibliography…
Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis.
The measurement of power spectra dover publications
RB Blackman and JW Tukey · 1958
Earlier work this paper cites.
Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences
Steven Davis and Paul Mermelstein · 1980
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Beat tracking by dynamic programming
Daniel PW Ellis · 2007
Earlier work this paper cites.
Cyclic tempogram—a mid-level tempo representation for musicsignals
Peter Grosche, Meinard Müller, and Frank Kurth · 2010
Earlier work this paper cites.
Constant-q transform toolbox for music processing
Christian Schörkhuber and Anssi Klapuri · 2010
Earlier work this paper cites.
Maximum filter vibrato suppression for onset detection
Sebastian Böck and Gerhard Widmer · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
librosa: Audio and music signal analysis in python
Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto · 2015
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Phase-functioned neural networks for character control
Daniel Holden, Taku Komura, and Jun Saito · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Dance with melody: An lstm-autoencoder approach to music-oriented dance synthesis
Taoran Tang, Jia Jia, and Hanyang Mao · 2018
Earlier work this paper cites.
Mode-adaptive neural networks for quadruped motion control
He Zhang, Sebastian Starke, Taku Komura, and Jun Saito · 2018
Earlier work this paper cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Earlier work this paper cites.
Self-supervised moving vehicle tracking with stereo sound
Chuang Gan, Hang Zhao, Peihao Chen, David Cox, and Antonio Torralba · 2019
Earlier work this paper cites.
Learning individual styles of conversational gesture
Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik · 2019
Earlier work this paper cites.
Expressive body capture: 3d hands, face, and body from a single image
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black · 2019
Earlier work this paper cites.
Diverse trajectory forecasting with determinantal point processes
Ye Yuan and Kris Kitani · 2019
Earlier work this paper cites.
On the continuity of rotation representations in neural networks
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li · 2019
Earlier work this paper cites.
Action2motion: Conditioned generation of 3d human motions
Chuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, and Li Cheng · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Speech gesture generation from the trimodal context of text, audio, and speaker identity
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee · 2020
Earlier work this paper cites.
Stochastic scene-aware motion prediction
Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael J Black · 2021
Earlier work this paper cites.
Ai choreographer: Music conditioned 3d dance generation with aist++
Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa · 2021
Earlier work this paper cites.
Speech drives templates: Co-speech gesture synthesis with learned templates
Shenhan Qian, Zhi Tu, Yihao Zhi, Wen Liu, and Shenghua Gao · 2021
Earlier work this paper cites.
Scene-aware generative network for human motion synthesis
Jingbo Wang, Sijie Yan, Bo Dai, and Dahua Lin · 2021
Earlier work this paper cites.
Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings
Tenglong Ao, Qingzhe Gao, Yuke Lou, Baoquan Chen, and Libin Liu · 2022
Earlier work this paper cites.
A bi-directional attention guided cross-modal network for music based dance generation
Di Fan, Lili Wan, Wanru Xu, and Shenghui Wang · 2022
Earlier work this paper cites.
Generating diverse and natural 3d human motions from text
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng · 2022
Earlier work this paper cites.
Danceformer: Music conditioned 3d dance generation with parametric motion transformer
Buyu Li, Yongchi Zhao, Shi Zhelun, and Lu Sheng · 2022
Earlier work this paper cites.
Learning hierarchical cross-modal association for co-speech gesture generation
Xian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu, Rui Qian, Xinyi Lin, Xiaowei Zhou, Wayne Wu, Bo Dai, and Bolei Zhou · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Earlier work this paper cites.
Bailando: 3d dance generation by actor-critic gpt with choreographic memory
Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu · 2022
Earlier work this paper cites.
Deepphase: Periodic autoencoders for learning motion phase manifolds
Sebastian Starke, Ian Mason, and Taku Komura · 2022
Earlier work this paper cites.
Towards diverse and natural scene-aware 3d human motion synthesis
Jingbo Wang, Yu Rong, Jingyuan Liu, Sijie Yan, Dahua Lin, and Bo Dai · 2022
Cited alongside, same era.
Music2dance: Dancenet for music-driven dance generation
Wenlin Zhuang, Congyi Wang, Jinxiang Chai, Yangang Wang, Ming Shao, and Siyu Xia · 2022
Cited alongside, same era.
Listen, denoise, action! audio-driven motion synthesis with diffusion models
Simon Alexanderson, Rajmund Nagy, Jonas Beskow, and Gustav Eje Henter · 2023
Cited alongside, same era.
Gesturediffuclip: Gesture diffusion model with clip latents
Tenglong Ao, Zeyi Zhang, and Libin Liu · 2023
Cited alongside, same era.
Belfusion: Latent diffusion for behavior-driven human motion prediction
German Barquero, Sergio Escalera, and Cristina Palmero · 2023
Cited alongside, same era.
Humanmac: Masked motion completion for human motion prediction
Ling-Hao Chen, Jiawei Zhang, Yewen Li, Yiren Pang, Xiaobo Xia, and Tongliang Liu · 2023
Modeling and driving human body soundfields through acoustic primitives
Chao Huang, Dejan Marković, Chenliang Xu, and Alexander Richard · 2024
Later among the works it cites.
Como: Controllable motion generation through language guided pose code editing
Yiming Huang, Weilin Wan, Yue Yang, Chris Callison-Burch, Mark Yatskar, and Lingjie Liu · 2024
Later among the works it cites.
Yinghao Huang, Leo Ho, Dafei Qin, Mingyi Shi, and Taku Komura · 2024
Later among the works it cites.
Acoustic volume rendering for neural impulse response fields
Zitong Lan, Chenhao Zheng, Zhiwei Zheng, and Mingmin Zhao · 2024
Later among the works it cites.
Lodge: A coarse to fine diffusion network for long dance generation guided by the characteristic dance primitives
Ronghui Li, YuXiang Zhang, Yachao Zhang, Hongwen Zhang, Jie Guo, Yan Zhang, Yebin Liu, and Xiu Li · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Executing your commands via motion diffusion in latent space
Xin Chen, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, and Gang Yu · 2023
Cited alongside, same era.
Mofusion: A framework for denoising-diffusion-based motion synthesis
Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, and Christian Theobalt · 2023
Cited alongside, same era.
C· ase: Learning conditional adversarial skill embeddings for physics-based characters
Zhiyang Dou, Xuelin Chen, Qingnan Fan, Taku Komura, and Wenping Wang · 2023
Cited alongside, same era.
Tm2d: Bimodality driven 3d dance generation via music-text integration
Kehong Gong, Dongze Lian, Heng Chang, Chuan Guo, Zihang Jiang, Xinxin Zuo, Michael Bi Mi, and Xinchao Wang · 2023
Cited alongside, same era.
Motiongpt: Human motion as a foreign language
Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen · 2023
Cited alongside, same era.
Controllable group choreography using contrastive diffusion
Nhat Le, Tuong Do, Khoa Do, Hien Nguyen, Erman Tjiputra, Quang D Tran, and Anh Nguyen · 2023
Cited alongside, same era.
Later among the works it cites.
Zhouyingcheng Liao, Mingyuan Zhang, Wenjia Wang, Lei Yang, and Taku Komura · 2024
Later among the works it cites.
Scamo: Exploring the scaling law in autoregressive motion generation model
Shunlin Lu, Jingbo Wang, Zeyu Lu, Ling-Hao Chen, Wenxun Dai, Junting Dong, Zhiyang Dou, Bo Dai, and Ruimao Zhang · 2024
Later among the works it cites.
Popdg: Popular 3d dance generation with popdanceset
Zhenye Luo, Min Ren, Xuecai Hu, Yongzhen Huang, and Li Yao · 2024
Later among the works it cites.
Mingyi Shi, Dafei Qin, Leo Ho, Zhouyingcheng Liao, Yinghao Huang, Junichi Yamagishi, and Taku Komura · 2024
Later among the works it cites.
Interactive character control with auto-regressive motion diffusion models
Yi Shi, Jingbo Wang, Xuekun Jiang, Bingkun Lin, Bo Dai, and Xue Bin Peng · 2024
Later among the works it cites.
Duolando: Follower gpt with off-policy reinforcement learning for dance accompaniment
Li Siyao, Tianpei Gu, Zhitao Yang, Zhengyu Lin, Ziwei Liu, Henghui Ding, Lei Yang, and Chen Change Loy · 2024
Later among the works it cites.
Comusion: Towards consistent stochastic human motion prediction via motion diffusion
Jiarui Sun and Girish Chowdhary · 2024
Later among the works it cites.
Both ears wide open: Towards language-driven spatial audio generation
Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye, Huadai Liu, Honggang Zhang, Wei Xue, and Yike Guo · 2024
Later among the works it cites.
Maskedmimic: Unified physics-based character control through masked motion inpainting
Chen Tessler, Yunrong Guo, Ofir Nabati, Gal Chechik, and Xue Bin Peng · 2024
Later among the works it cites.
Sims: Simulating stylized human-scene interactions with retrieval-augmented script generation
Wenjia Wang, Liang Pan, Zhiyang Dou, Jidong Mei, Zhouyingcheng Liao, Yuke Lou, Yifan Wu, Lei Yang, Jingbo Wang, and Taku Komura · 2024
Later among the works it cites.
Move as you say interact as you can: Language-guided human motion generation with scene affordance
Zan Wang, Yixin Chen, Baoxiong Jia, Puhao Li, Jinlu Zhang, Jingze Zhang, Tengyu Liu, Yixin Zhu, Wei Liang, and Siyuan Huang · 2024
Later among the works it cites.
Unimumo: Unified text, music and motion generation
Han Yang, Kun Su, Yutong Zhang, Jiaben Chen, Kaizhi Qian, Gaowen Liu, and Chuang Gan · 2024
Later among the works it cites.
Bidirectional autoregessive diffusion model for dance generation
Canyu Zhang, Youbao Tang, Ning Zhang, Ruei-Sung Lin, Mei Han, Jing Xiao, and Song Wang · 2024
Later among the works it cites.
Large motion model for unified multi-modal motion generation
Mingyuan Zhang, Daisheng Jin, Chenyang Gu, Fangzhou Hong, Zhongang Cai, Jingfang Huang, Chongzhi Zhang, Xinying Guo, Lei Yang, Ying He, et al · 2024
Later among the works it cites.
Physpt: Physics-aware pretrained transformer for estimating human dynamics from monocular videos
Yufei Zhang, Jeffrey O Kephart, Zijun Cui, and Qiang Ji · 2024
Later among the works it cites.
Incorporating physics principles for precise human motion prediction
Yufei Zhang, Jeffrey O Kephart, and Qiang Ji · 2024
Later among the works it cites.
Semantic gesticulator: Semantics-aware co-speech gesture synthesis
Zeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao, Chuan Lin, Baoquan Chen, and Libin Liu · 2024
Later among the works it cites.
Bat: Learning to reason about spatial sounds with large language models
Zhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen, Eunsol Choi, and David Harwath · 2024
Later among the works it cites.
Motioncraft: Crafting whole-body motion with plug-and-play multimodal controls
Yuxuan Bian, Ailing Zeng, Xuan Ju, Xian Liu, Zhaoyang Zhang, Wei Liu, and Qiang Xu · 2025
Closest in time.
Go to zero: Towards zero-shot motion generation with million-scale data
Ke Fan, Shunlin Lu, Minyue Dai, Runyi Yu, Lixing Xiao, Zhiyang Dou, Junting Dong, Lizhuang Ma, and Jingbo Wang · 2025
Closest in time.
Asap: Aligning simulation and real-world physics for learning agile humanoid whole-body skills
Tairan He, Jiawei Gao, Wenli Xiao, Yuanhang Zhang, Zi Wang, Jiashun Wang, Zhengyi Luo, Guanqi He, Nikhil Sobanbab, Chaoyi Pan, et al · 2025
Closest in time.
Modskill: Physical character skill modularization
Yiming Huang, Zhiyang Dou, and Lingjie Liu · 2025
Closest in time.
Visage: Video-to-spatial audio generation
Jaeyeon Kim, Heeseung Yun, and Gunhee Kim · 2025
Closest in time.
Resounding acoustic fields with reciprocity, 2025
Zitong Lan, Yiduo Hao, and Mingmin Zhao · 2025
Closest in time.
Vicon motion capture systems, 2025
Vicon Motion Systems Ltd · 2025
Closest in time.
Tokenhsi: Unified synthesis of physical human-scene interactions through task tokenization
Liang Pan, Zeshi Yang, Zhiyang Dou, Wenjia Wang, Buzhen Huang, Bo Dai, Taku Komura, and Jingbo Wang · 2025
Closest in time.
Coda: Coordinated diffusion noise optimization for whole-body manipulation of articulated objects
Huaijin Pi, Zhi Cen, Zhiyang Dou, and Taku Komura · 2025
Closest in time.
Intermimic: Towards universal whole-body control for physics-based human-object interactions
Sirui Xu, Hung Yu Ling, Yu-Xiong Wang, and Liang-Yan Gui · 2025
Closest in time.
Emdm: Efficient motion diffusion model for fast and high-quality motion generation
Wenyang Zhou, Zhiyang Dou, Zeyu Cao, Zhouyingcheng Liao, Jingbo Wang, Wenjia Wang, Yuan Liu, Taku Komura, Wenping Wang, and Lingjie Liu · 2025
Closest in time.