Fetching the paper…
Reading the bibliography…
Recent work has demonstrated the significant potential of denoising diffusion models for generating human motion, including text-to-motion capabilities.
Motion graphs
Lucas Kovar, Michael Gleicher, and Frédéric Pighin. 2008 · 2008
Earlier work this paper cites.
Panoptic Studio: A Massively Multiview System for Social Motion Capture. In The IEEE International Conference on Computer Vision (ICCV)
Hanbyul Joo, Hao Liu, Lei Tan, Lin Gui, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, and Yaser Sheikh. 2015 · 2015
Earlier work this paper cites.
SMPL: A skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning . PMLR, 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
On human motion prediction using recurrent neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2891–2900
Julieta Martinez, Michael J Black, and Javier Romero. 2017 · 2017
Earlier work this paper cites.
Single-shot multi-person 3d pose estimation from monocular rgb. In 2018 International Conference on 3D Vision (3DV) . IEEE, 120–130
Dushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu, Srinath Sridhar, Gerard Pons-Moll, and Christian Theobalt. 2018 · 2018
Earlier work this paper cites.
Recovering accurate 3d human pose in the wild using imus and a moving camera. In Proceedings of the European Conference on Computer Vision (ECCV) . 601–617
Timo Von Marcard, Roberto Henschel, Michael J Black, Bodo Rosenhahn, and Gerard Pons-Moll. 2018 · 2018
Earlier work this paper cites.
A sampling approach to generating closely interacting 3d pose-pairs from 2d annotations
Kangxue Yin, Hui Huang, Edmond SL Ho, Hao Wang, Taku Komura, Daniel Cohen-Or, and Hao Zhang. 2018 · 2018
Earlier work this paper cites.
Auto-Conditioned Recurrent Networks for Extended Complex Human Motion Synthesis. In International Conference on Learning Representations
Yi Zhou, Zimo Li, Shuangjiu Xiao, Chong He, Zeng Huang, and Hao Li. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 4401–4410
Tero Karras, Samuli Laine, and Timo Aila. 2019 · 2019
Earlier work this paper cites.
AMASS: Archive of Motion Capture as Surface Shapes. In International Conference on Computer Vision . 5442–5451
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. 2019 · 2019
Earlier work this paper cites.
Expressive Body Capture: 3D Hands, Face, and Body from a Single Image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) . 10975–10985
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. 2019 · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2020 · 2020
Earlier work this paper cites.
ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . 14347–14356
Jooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon, and Sungroh Yoon. 2021 · 2021
Cited alongside, same era.
BABEL: Bodies, Action and Behavior with English Labels. In Proceedings IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) . 722–731
Abhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J. Black. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision. In International Conference on Machine Learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Multi-Person 3D Motion Prediction with Multi-Range Transformers
Jiashun Wang, Huazhe Xu, Medhini Narasimhan, and Xiaolong Wang. 2021 · 2021
Cited alongside, same era.
StyleAlign: Analysis and Applications of Aligned StyleGAN Models
MoDi: Unconditional Motion Synthesis from Diverse Data
Sigal Raab, Inbal Leibovitch, Peizhuo Li, Kfir Aberman, Olga Sorkine-Hornung, and Daniel Cohen-Or. 2022 · 2022
Later among the works it cites.
Palette: Image-to-image diffusion models. In ACM SIGGRAPH 2022 Conference Proceedings . 1–10
Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi. 2022 · 2022
Later among the works it cites.
ActFormer: A GAN Transformer Framework towards General Action-Conditioned 3D Human Motion Generation
Ziyang Song, Dongliang Wang, Nan Jiang, Zhicheng Fang, Chenjing Ding, Weihao Gan, and Wei Wu. 2022 · 2022
Later among the works it cites.
Motionclip: Exposing human motion generation to clip space. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXII . Springer, 358–374
Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zongze Wu, Yotam Nitzan, Eli Shechtman, and Dani Lischinski. 2021 · 2021
Cited alongside, same era.
TEACH: Temporal Action Compositions for 3D Humans. In International Conference on 3D Vision (3DV)
Nikos Athanasiou, Mathis Petrovich, Michael J. Black, and Gül Varol. 2022 · 2022
Cited alongside, same era.
Generating Diverse and Natural 3D Human Motions From Text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5152–5161
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. 2022 · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. 2022 · 2022
Cited alongside, same era.
AvatarCLIP: Zero-Shot Text-Driven Generation and Animation of 3D Avatars
Fangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai, Lei Yang, and Ziwei Liu. 2022 · 2022
Cited alongside, same era.
FLAME: Free-form Language-based Motion Synthesis & Editing
Jihoon Kim, Jiseob Kim, and Sungjoon Choi. 2022 · 2022
Cited alongside, same era.
Repaint: Inpainting using denoising diffusion probabilistic models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11461–11471
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. 2022 · 2022
Cited alongside, same era.
Weakly-supervised Action Transition Learning for Stochastic Human Motion Prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8151–8160
Wei Mao, Miaomiao Liu, and Mathieu Salzmann. 2022 · 2022
Cited alongside, same era.
Pose-ndf: Modeling human pose manifolds with neural distance fields. In European Conference on Computer Vision . Springer, 572–589
Garvita Tiwari, Dimitrije Antić, Jan Eric Lenssen, Nikolaos Sarafianos, Tony Tung, and Gerard Pons-Moll. 2022 · 2022
Later among the works it cites.
EDGE: Editable Dance Generation From Music
Jonathan Tseng, Rodrigo Castellon, and C Karen Liu. 2022 · 2022
Later among the works it cites.
SoMoFormer: Multi-Person Pose Forecasting with Transformers
Edward Vendrow, Satyajit Kumar, Ehsan Adeli, and Hamid Rezatofighi. 2022 · 2022
Later among the works it cites.
NEURAL MARIONETTE: A Transformer-based Multi-action Human Motion Synthesis System
Weiqiang Wang, Xuefei Zhe, Huan Chen, Di Kang, Tingguang Li, Ruizhi Chen, and Linchao Bao. 2022 · 2022
Later among the works it cites.
Executing your Commands via Motion Diffusion in Latent Space
Chen Xin, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, Jingyi Yu, and Gang Yu. 2022 · 2022
Later among the works it cites.
PhysDiff: Physics-Guided Human Motion Diffusion Model
Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz. 2022 · 2022
Later among the works it cites.
MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. 2022 · 2022
Later among the works it cites.
Mofusion: A framework for denoising-diffusion-based motion synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9760–9770
Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, and Christian Theobalt. 2023 · 2023
Closest in time.
Sigal Raab, Inbal Leibovitch, Guy Tevet, Moab Arar, Amit H Bermano, and Daniel Cohen-Or. 2023 · 2023
Closest in time.
Human Motion Diffusion Model. In The Eleventh International Conference on Learning Representations
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. 2023 · 2023
Closest in time.