Fetching the paper…
Reading the bibliography…
Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation.
Locke: Epistemology and ontology
Michael Ayers · 1991
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Aapo Hyvärinen and Peter Dayan · 2005
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, Fujie Huang, et al · 2006
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Improved contrastive divergence training of energy based models
Yilun Du, Shuang Li, Joshua Tenenbaum, and Igor Mordatch · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
The kit motion-language dataset
Matthias Plappert, Christian Mandery, and Tamim Asfour · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Energy-based generative adversarial networks
Junbo Zhao, Michael Mathieu, and Yann LeCun · 2017
Earlier work this paper cites.
Language2pose: Natural language grounded pose forecasting
Chaitanya Ahuja and Louis-Philippe Morency · 2019
Earlier work this paper cites.
Implicit generation and modeling with energy based models
Yilun Du and Igor Mordatch · 2019
Earlier work this paper cites.
Amass: Archive of motion capture as surface shapes
Naureen Mahmood, Nima Ghorbani, Nikolaus F Troje, Gerard Pons-Moll, and Michael J Black · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Hopfield networks is all you need
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, et al · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2020
Earlier work this paper cites.
Unsupervised learning of compositional energy concepts
Yilun Du, Shuang Li, Yash Sharma, Josh Tenenbaum, and Igor Mordatch · 2021
Earlier work this paper cites.
Synthesis of compositional animations from textual descriptions
Anindita Ghosh, Noshaba Cheema, Cennet Oguz, Christian Theobalt, and Philipp Slusallek · 2021
Earlier work this paper cites.
Learning to compose visual relations
Nan Liu, Shuang Li, Yilun Du, Josh Tenenbaum, and Antonio Torralba · 2021
Earlier work this paper cites.
Controllable and compositional generation with latent-space energy-based models
Weili Nie, Arash Vahdat, and Anima Anandkumar · 2021
Earlier work this paper cites.
Action-conditioned 3D human motion synthesis with transformer VAE
Mathis Petrovich, Michael J. Black, and Gül Varol · 2021
Earlier work this paper cites.
TEACH: Temporal Action Compositions for 3D Humans
Nikos Athanasiou, Mathis Petrovich, Michael J. Black, and Gül Varol · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Cited alongside, same era.
Stylet2i: Toward compositional and high-fidelity text-to-image synthesis
Zhiheng Li, Martin Renqiang Min, Kai Li, and Chenliang Xu · 2022
Cited alongside, same era.
Compositional visual generation with composable diffusion models
Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum · 2022
Cited alongside, same era.
TEMOS: Generating diverse human motions from textual descriptions
Mathis Petrovich, Michael J. Black, and Gul Varol · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis
Mathis Petrovich, Michael J Black, and Gül Varol · 2023
Later among the works it cites.
Realistic human motion generation with cross-diffusion models
Zeping Ren, Shaoli Huang, and Xiu Li · 2023
Later among the works it cites.
Exploring compositional visual generation with latent classifier guidance
Changhao Shi, Haomiao Ni, Kai Li, Shaobo Han, Mingfu Liang, and Martin Renqiang Min · 2023
Later among the works it cites.
Fg-t2m: Fine-grained text-driven human motion generation via diffusion model
Yin Wang, Zhiying Leng, Frederick WB Li, Shun-Cheng Wu, and Xiaohui Liang · 2023
Later among the works it cites.
Omnicontrol: Control any joint at any time for human motion generation
Yiming Xie, Varun Jampani, Lei Zhong, Deqing Sun, and Huaizu Jiang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jonathan Tseng, Rodrigo Castellon, and C Karen Liu · 2022
Cited alongside, same era.
Executing your commands via motion diffusion in latent space
Chen Xin, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, Jingyi Yu, and Gang Yu · 2022
Cited alongside, same era.
Motiondiffuse: Text-driven human motion generation with diffusion model
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu · 2022
Cited alongside, same era.
SINC: Spatial composition of 3D human motions for simultaneous action generation
Nikos Athanasiou, Mathis Petrovich, Michael J. Black, and Gül Varol · 2023
Cited alongside, same era.
Attribute-centric compositional text-to-image generation
Yuren Cong, Martin Renqiang Min, Li Erran Li, Bodo Rosenhahn, and Michael Ying Yang · 2023
Cited alongside, same era.
Mofusion: A framework for denoising-diffusion-based motion synthesis
Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, and Christian Theobalt · 2023
Cited alongside, same era.
Reduce, reuse, recycle: Compositional generation with energy-based diffusion models and mcmc
Yilun Du, Conor Durkan, Robin Strudel, Joshua B Tenenbaum, Sander Dieleman, Rob Fergus, Jascha Sohl-Dickstein, Arnaud Doucet, and Will Sussman Grathwohl · 2023
Cited alongside, same era.
Synthesizing long-term human motions with diffusion models via coherent sampling
Zhao Yang, Bing Su, and Ji-Rong Wen · 2023
Later among the works it cites.
Physdiff: Physics-guided human motion diffusion model
Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz · 2023
Later among the works it cites.
Language-guided human motion synthesis with atomic actions
Yuanhao Zhai, Mingzhen Huang, Tianyu Luan, Lu Dong, Ifeoma Nwogu, Siwei Lyu, David Doermann, and Junsong Yuan · 2023
Later among the works it cites.
Attt2m: Text-driven human motion generation with multi-perspective attention mechanism
Chongyang Zhong, Lei Hu, Zihao Zhang, and Shihong Xia · 2023
Later among the works it cites.
Ude: A unified driving engine for human motion generation
Zixiang Zhou and Baoyuan Wang · 2023
Later among the works it cites.
Iterative motion editing with natural language
Purvi Goel, Kuan-Chieh Wang, C Karen Liu, and Kayvon Fatahalian · 2024
Closest in time.
Momask: Generative masked modeling of 3d human motions
Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang, and Li Cheng · 2024
Closest in time.
Amd: Autoregressive motion diffusion
Bo Han, Hao Peng, Minjing Dong, Yi Ren, Yixuan Shen, and Chang Xu · 2024
Closest in time.
Energy transformer
Benjamin Hoover, Yuchen Liang, Bao Pham, Rameswar Panda, Hendrik Strobelt, Duen Horng Chau, Mohammed Zaki, and Dmitry Krotov · 2024
Closest in time.
Vision-language navigation with energy-based policy
Rui Liu, Wenguan Wang, and Yi Yang · 2024
Closest in time.
Energy-based cross attention for bayesian context update in text-to-image diffusion models
Geon Yeong Park, Jeongsol Kim, Beomsu Kim, Sang Wan Lee, and Jong Chul Ye · 2024
Closest in time.
Multi-track timeline control for text-driven 3d human motion generation
Mathis Petrovich, Or Litany, Umar Iqbal, Michael J. Black, Gül Varol, Xue Bin Peng, and Davis Rempe · 2024
Closest in time.
Mmm: Generative masked motion model
Ekkasit Pinyoanuntapong, Pu Wang, Minwoo Lee, and Chen Chen · 2024
Closest in time.
Human motion diffusion as a generative prior
Yoni Shafir, Guy Tevet, Roy Kapon, and Amit Haim Bermano · 2024
Closest in time.
Towards detailed text-to-motion synthesis via basic-to-advanced hierarchical diffusion model
Zhenyu Xie, Yang Wu, Xuehao Gao, Zhongqian Sun, Wei Yang, and Xiaodan Liang · 2024
Closest in time.
Motion mamba: Efficient and long sequence motion generation
Zeyu Zhang, Akide Liu, Ian Reid, Richard Hartley, Bohan Zhuang, and Hao Tang · 2025
Closest in time.