Fetching the paper…
Reading the bibliography…
State-of-the-art text-to-motion generation models rely on the kinematic-aware, local-relative motion representation popularized by HumanML3D, which encodes motion relative to the pelvis and to the previous frame with built-in redundancy.
Learning internal representations by error propagation
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
Auto-encoding variational bayes, 2013
Diederik P Kingma, Max Welling, et al · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Smpl: a skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black · 2015
Earlier work this paper cites.
Keep it smpl: Automatic estimation of 3d human pose and shape from a single image
Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J Black · 2016
Earlier work this paper cites.
The kit motion-language dataset
Matthias Plappert, Christian Mandery, and Tamim Asfour · 2016
Earlier work this paper cites.
Carnegie mellon university - cmu graphics lab - motion capture library
Carnegie Mellon University CMU Graphics Lab motion capture library · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Generating animated videos of human activities from natural language descriptions
Angela S. Lin, Lemeng Wu, Rodolfo Corona, Kevin W. H. Tai, Qi-Xing Huang, and Raymond J. Mooney · 2018
Earlier work this paper cites.
Paired recurrent autoencoders for bidirectional translation between robot actions and linguistic descriptions
Tatsuro Yamada, Hiroyuki Matsunaga, and Tetsuya Ogata · 2018
Earlier work this paper cites.
Language2pose: Natural language grounded pose forecasting
Chaitanya Ahuja and Louis-Philippe Morency · 2019
Earlier work this paper cites.
Amass: Archive of motion capture as surface shapes
Naureen Mahmood, Nima Ghorbani, Nikolaus F Troje, Gerard Pons-Moll, and Michael J Black · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Action2motion: Conditioned generation of 3d human motions
Chuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, and Li Cheng · 2020
Earlier work this paper cites.
Query-key normalization for transformers
Alex Henry, Prudhvi Raj Dachapally, Shubham Pawar, and Yuxuan Chen · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Earlier work this paper cites.
Fully convolutional mesh autoencoder using efficient spatially varying kernels
Yi Zhou, Chenglei Wu, Zimo Li, Chen Cao, Yuting Ye, Jason Saragih, Hao Li, and Yaser Sheikh · 2020
Earlier work this paper cites.
Text2gestures: A transformer-based network for generating emotive body gestures for virtual agents
Uttaran Bhattacharya, Nicholas Rewkowski, Abhishek Banerjee, Pooja Guhan, Aniket Bera, and Dinesh Manocha · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Earlier work this paper cites.
Synthesis of compositional animations from textual descriptions
Anindita Ghosh, Noshaba Cheema, Cennet Oguz, Christian Theobalt, and Philipp Slusallek · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Maskgit: Masked generative image transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman · 2022
Earlier work this paper cites.
Generating diverse and natural 3d human motions from text
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng · 2022
Earlier work this paper cites.
Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts
Chuan Guo, Xinxin Zuo, Sen Wang, and Li Cheng · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Earlier work this paper cites.
Flow straight and fast: Learning to generate and transfer data with rectified flow
Xingchao Liu, Chengyue Gong, and Qiang Liu · 2022
Earlier work this paper cites.
Temos: Generating diverse human motions from textual descriptions
Mathis Petrovich, Michael J Black, and Gül Varol · 2022
Earlier work this paper cites.
Motionclip: Exposing human motion generation to clip space
Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or · 2022
Earlier work this paper cites.
Towards diverse and natural scene-aware 3d human motion synthesis
Jingbo Wang, Yu Rong, Jingyuan Liu, Sijie Yan, Dahua Lin, and Bo Dai · 2022
Earlier work this paper cites.
Motiondiffuse: Text-driven human motion generation with diffusion model
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu · 2022
Earlier work this paper cites.
Make-an-animation: Large-scale text-conditional 3d human motion generation
Samaneh Azadi, Akbar Shah, Thomas Hayes, Devi Parikh, and Sonal Gupta · 2023
Earlier work this paper cites.
Autoencoders
Dor Bank, Noam Koenigstein, and Raja Giryes · 2023
Earlier work this paper cites.
Executing your commands via motion diffusion in latent space
Xin Chen, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, and Gang Yu · 2023
Earlier work this paper cites.
Remos: Reactive 3d motion synthesis for two-person interactions
Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik, Christian Theobalt, and Philipp Slusallek · 2023
Earlier work this paper cites.
Motion flow matching for human motion synthesis and editing
Vincent Tao Hu, Wenzhe Yin, Pingchuan Ma, Yunlu Chen, Basura Fernando, Yuki M Asano, Efstratios Gavves, Pascal Mettes, Bjorn Ommer, and Cees GM Snoek · 2023
Earlier work this paper cites.
Diffusion-based generation, optimization, and planning in 3d scenes
Siyuan Huang, Zan Wang, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu, Wei Liang, and Song-Chun Zhu · 2023
Earlier work this paper cites.
Motiongpt: Human motion as a foreign language
Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen · 2023
Earlier work this paper cites.
Guided motion diffusion for controllable human motion synthesis
Korrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn, and Siyu Tang · 2023
Earlier work this paper cites.
Flame: Free-form language-based motion synthesis & editing
Jihoon Kim, Jiseob Kim, and Sungjoon Choi · 2023
Earlier work this paper cites.
Object motion guided human motion synthesis
Jiaman Li, Jiajun Wu, and C Karen Liu · 2023
Earlier work this paper cites.
Motion-x: A large-scale 3d expressive whole-body human motion dataset
Jing Lin, Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang · 2023
Earlier work this paper cites.
Diversemotion: Towards diverse human motion generation via discrete diffusion
Yunhong Lou, Linchao Zhu, Yaxiong Wang, Xiaohan Wang, and Yi Yang · 2023
Earlier work this paper cites.
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Earlier work this paper cites.
Hoi-diff: Text-driven synthesis of 3d human-object interactions using diffusion models
Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani, Deqing Sun, and Huaizu Jiang · 2023
Cited alongside, same era.
TMR: Text-to-motion retrieval using contrastive 3D human motion synthesis
Mathis Petrovich, Michael J. Black, and Gül Varol · 2023
Cited alongside, same era.
Hierarchical generation of human-object interactions with diffusion probabilistic models
Huaijin Pi, Sida Peng, Minghui Yang, Xiaowei Zhou, and Hujun Bao · 2023
Cited alongside, same era.
Trace and pace: Controllable pedestrian animation via guided trajectory diffusion
Davis Rempe, Zhengyi Luo, Xue Bin Peng, Ye Yuan, Kris Kitani, Karsten Kreis, Sanja Fidler, and Or Litany · 2023
Cited alongside, same era.
Human motion diffusion as a generative prior
Yonatan Shafir, Guy Tevet, Roy Kapon, and Amit H Bermano · 2023
Contact-aware human motion generation from textual descriptions
Sihan Ma, Qiong Cao, Jing Zhang, and Dacheng Tao · 2024
Later among the works it cites.
Rethinking diffusion for text-driven human motion generation
Zichong Meng, Yiming Xie, Xiaogang Peng, Zeyu Han, and Huaizu Jiang · 2024
Later among the works it cites.
Multi-track timeline control for text-driven 3d human motion generation
Mathis Petrovich, Or Litany, Umar Iqbal, Michael J Black, Gul Varol, Xue Bin Peng, and Davis Rempe · 2024
Later among the works it cites.
Motion-2-to-3: Leveraging 2d motion data to boost 3d motion generation
Huaijin Pi, Ruoxi Guo, Zehong Shen, Qing Shuai, Zechen Hu, Zhumei Wang, Yajiao Dong, Ruizhen Hu, Taku Komura, Sida Peng, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generating fine-grained human motions using chatgpt-refined descriptions
Xu Shi, Chuanchen Luo, Junran Peng, Hongwen Zhang, and Yunlian Sun · 2023
Cited alongside, same era.
Human motion diffusion model
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano · 2023
Cited alongside, same era.
Tlcontrol: Trajectory and language control for human motion synthesis
Weilin Wan, Zhiyang Dou, Taku Komura, Wenping Wang, Dinesh Jayaraman, and Lingjie Liu · 2023
Cited alongside, same era.
Diffusionphase: Motion diffusion in frequency domain
Weilin Wan, Yiming Huang, Shutong Wu, Taku Komura, Wenping Wang, Dinesh Jayaraman, and Lingjie Liu · 2023
Cited alongside, same era.
Intercontrol: Generate human motion interactions by controlling every joint
Zhenzhi Wang, Jingbo Wang, Dahua Lin, and Bo Dai · 2023
Cited alongside, same era.
Interdiff: Generating 3d human-object interactions with physics-informed diffusion
Sirui Xu, Zhengyuan Li, Yu-Xiong Wang, and Liang-Yan Gui · 2023
Cited alongside, same era.
Cross-modal retrieval for motion and text via droptriple loss
Sheng Yan, Yang Liu, Haoqiang Wang, Xin Du, Mengyuan Liu, and Hong Liu · 2023
Cited alongside, same era.
Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Korrawe Karunratanakul, Pu Wang, Hongfei Xue, Chen Chen, Chuan Guo, Junli Cao, Jian Ren, and Sergey Tulyakov · 2024
Later among the works it cites.
Bamm: Bidirectional autoregressive motion model
Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Pu Wang, Minwoo Lee, Srijan Das, and Chen Chen · 2024
Later among the works it cites.
Mmm: Generative masked motion model
Ekkasit Pinyoanuntapong, Pu Wang, Minwoo Lee, and Chen Chen · 2024
Later among the works it cites.
Interactive character control with auto-regressive motion diffusion models
Yi Shi, Jingbo Wang, Xuekun Jiang, Bingkun Lin, Bo Dai, and Xue Bin Peng · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Later among the works it cites.
Coma: Compositional human motion generation with multi-modal agents
Shanlin Sun, Gabriel De Araujo, Jiaqi Xu, Shenghan Zhou, Hanwen Zhang, Ziheng Huang, Chenyu You, and Xiaohui Xie · 2024
Later among the works it cites.
Closd: Closing the loop between simulation and diffusion for multi-task character control
Guy Tevet, Sigal Raab, Setareh Cohan, Daniele Reda, Zhengyi Luo, Xue Bin Peng, Amit H Bermano, and Michiel van de Panne · 2024
Later among the works it cites.
Sims: Simulating human-scene interactions with real world script planning
Wenjia Wang, Liang Pan, Zhiyang Dou, Zhouyingcheng Liao, Yuke Lou, Lei Yang, Jingbo Wang, and Taku Komura · 2024
Later among the works it cites.
Text-controlled motion mamba: Text-instructed temporal grounding of human motion
Xinghan Wang, Zixi Kang, and Yadong Mu · 2024
Later among the works it cites.
Thor: Text to human-object interaction diffusion via relation intervention
Qianyang Wu, Ye Shi, Xiaoshui Huang, Jingyi Yu, Lan Xu, and Jingya Wang · 2024
Later among the works it cites.
Omnicontrol: Control any joint at any time for human motion generation
Yiming Xie, Varun Jampani, Lei Zhong, Deqing Sun, and Huaizu Jiang · 2024
Later among the works it cites.
Motionbank: A large-scale video motion benchmark with disentangled rule-based annotations
Liang Xu, Shaoyang Hua, Zili Lin, Yifan Liu, Feipeng Ma, Yichao Yan, Xin Jin, Xiaokang Yang, and Wenjun Zeng · 2024
Later among the works it cites.
Inter-x: Towards versatile human-human interaction analysis
Liang Xu, Xintao Lv, Yichao Yan, Xin Jin, Shuwen Wu, Congsheng Xu, Yifan Liu, Yizhou Zhou, Fengyun Rao, Xingdong Sheng, et al · 2024
Later among the works it cites.
Interdreamer: Zero-shot text to 3d dynamic human-object interaction
Sirui Xu, Ziyin Wang, Yu-Xiong Wang, and Liang-Yan Gui · 2024
Later among the works it cites.
Mogents: Motion generation based on spatial-temporal joint modeling
Weihao Yuan, Weichao Shen, Yisheng He, Yuan Dong, Xiaodong Gu, Zilong Dong, Liefeng Bo, and Qixing Huang · 2024
Later among the works it cites.
Energymogen: Compositional human motion generation with energy-based diffusion model in latent space
Jianrong Zhang, Hehe Fan, and Yi Yang · 2024
Later among the works it cites.
Motiongpt: Finetuned llms are general-purpose motion generators
Yaqi Zhang, Di Huang, Bin Liu, Shixiang Tang, Yan Lu, Lu Chen, Lei Bai, Qi Chu, Nenghai Yu, and Wanli Ouyang · 2024
Later among the works it cites.
Motion mamba: Efficient and long sequence motion generation
Zeyu Zhang, Akide Liu, Ian Reid, Richard Hartley, Bohan Zhuang, and Hao Tang · 2024
Later among the works it cites.
Tedi: Temporally-entangled diffusion for long-term motion synthesis
Zihan Zhang, Richard Liu, Rana Hanocka, and Kfir Aberman · 2024
Later among the works it cites.
Dart: A diffusion-based autoregressive motion model for real-time text-driven motion control
Kaifeng Zhao, Gen Li, and Siyu Tang · 2024
Later among the works it cites.
Infinidreamer: Arbitrarily long human motion generation via segment score distillation
Wenjie Zhuo, Fan Ma, and Hehe Fan · 2024
Later among the works it cites.
Ready-to-react: Online reaction policy for two-character interaction generation
Zhi Cen, Huaijin Pi, Sida Peng, Qing Shuai, Yujun Shen, Hujun Bao, Xiaowei Zhou, and Ruizhen Hu · 2025
Closest in time.
Motionlcm: Real-time controllable motion generation via latent consistency model
Wenxun Dai, Ling-Hao Chen, Jingbo Wang, Jinpeng Liu, Bo Dai, and Yansong Tang · 2025
Closest in time.
Como: Controllable motion generation through language guided pose code editing
Yiming Huang, Weilin Wan, Yue Yang, Chris Callison-Burch, Mark Yatskar, and Lingjie Liu · 2025
Closest in time.
Scenemi: Motion in-betweening for modeling human-scene interactions
Inwoo Hwang, Bing Zhou, Young Min Kim, Jian Wang, and Chuan Guo · 2025
Closest in time.
Controllable human-object interaction synthesis
Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu, Xavier Puig, and C Karen Liu · 2025
Closest in time.
Shape my moves: Text-driven shape-aware synthesis of human motions
Ting-Hsuan Liao, Yi Zhou, Yu Shen, Chun-Hao Paul Huang, Saayan Mitra, Jia-Bin Huang, and Uttaran Bhattacharya · 2025
Closest in time.
Revisit human-scene interaction via space occupancy
Xinpeng Liu, Haowen Hou, Yanchao Yang, Yong-Lu Li, and Cewu Lu · 2025
Closest in time.
Zero-shot human-object interaction synthesis with multimodal priors
Yuke Lou, Yiming Wang, Zhen Wu, Rui Zhao, Wenjia Wang, Mingyi Shi, and Taku Komura · 2025
Closest in time.
Mixermdm: Learnable composition of human motion diffusion models
Pablo Ruiz-Ponce, German Barquero, Cristina Palmero, Sergio Escalera, and José García-Rodríguez · 2025
Closest in time.
Humos: Human motion model conditioned on body shape
Shashank Tripathi, Omid Taheri, Christoph Lassner, Michael Black, Daniel Holden, and Carsten Stoll · 2025
Closest in time.
Lixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan, Liang Pan, Yueer Zhou, Ziyong Feng, Xiaowei Zhou, Sida Peng, and Jingbo Wang · 2025
Closest in time.
Intermimic: Towards universal whole-body control for physics-based human-object interactions
Sirui Xu, Hung Yu Ling, Yu-Xiong Wang, and Liang-Yan Gui · 2025
Closest in time.
Guiding human-object interactions with rich geometry and relations
Mengqing Xue, Yifei Liu, Ling Guo, Shaoli Huang, and Changxing Ding · 2025
Closest in time.
Generating human interaction motions in scenes with text control
Hongwei Yi, Justus Thies, Michael J Black, Xue Bin Peng, and Davis Rempe · 2025
Closest in time.
Socialgen: Modeling multi-human social interaction with language models
Heng Yu, Juze Zhang, Changan Chen, Tiange Xiang, Yusu Fang, Juan Carlos Niebles, and Ehsan Adeli · 2025
Closest in time.
Motion anything: Any to motion generation
Zeyu Zhang, Yiran Wang, Wei Mao, Danning Li, Rui Zhao, Biao Wu, Zirui Song, Bohan Zhuang, Ian Reid, and Richard Hartley · 2025
Closest in time.
Smoodi: Stylized motion diffusion model
Lei Zhong, Yiming Xie, Varun Jampani, Deqing Sun, and Huaizu Jiang · 2025
Closest in time.
Emdm: Efficient motion diffusion model for fast and high-quality motion generation
Wenyang Zhou, Zhiyang Dou, Zeyu Cao, Zhouyingcheng Liao, Jingbo Wang, Wenjia Wang, Yuan Liu, Taku Komura, Wenping Wang, and Lingjie Liu · 2025
Closest in time.
Parco: Part-coordinating text-to-motion synthesis
Qiran Zou, Shangyuan Yuan, Shian Du, Yu Wang, Chang Liu, Yi Xu, Jie Chen, and Xiangyang Ji · 2025
Closest in time.