Fetching the paper…
Reading the bibliography…
This paper addresses the problem of generating 3D interactive human motion from text.
R. Hadsell, S. Chopra, and Y. LeCun, “Dimensionality reduction by learning an invariant mapping,” in 2006 IEEE computer society conference on computer vision and pattern recognition (CVPR’06) , vol. 2. IEEE, 2006, pp. 1735–1742
2006
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in International conference on machine learning . PMLR, 2015, pp. 2256–2265
2015
Earlier work this paper cites.
M. Plappert, C. Mandery, and T. Asfour, “The kit motion-language dataset,” Big data , vol. 4, no. 4, pp. 236–252, 2016
2016
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
Y. Ji, F. Xu, Y. Yang, F. Shen, H. T. Shen, and W.-S. Zheng, “A large-scale rgb-d database for arbitrary-view human action recognition,” in Proceedings of the 26th ACM international Conference on Multimedia , 2018, pp. 1510–1518
2018
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 5828–5839
2018
Earlier work this paper cites.
C. Ahuja and L.-P. Morency, “Language2pose: Natural language grounded pose forecasting,” in 2019 International Conference on 3D Vision (3DV) . IEEE, 2019, pp. 719–728
2019
Earlier work this paper cites.
J. Liu, A. Shahroudy, M. Perez, G. Wang, L.-Y. Duan, and A. C. Kot, “Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,” IEEE transactions on pattern analysis and machine intelligence , vol. 42, no. 10, pp. 2684–2701, 2019
2019
Earlier work this paper cites.
G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black, “Expressive body capture: 3d hands, face, and body from a single image,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 10 975–10 985
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. Rempe, L. J. Guibas, A. Hertzmann, B. Russell, R. Villegas, and J. Yang, “Contact and human dynamics from monocular video,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part V 16 . Springer, 2020, pp. 71–87
2020
Earlier work this paper cites.
S. Shimada, V. Golyanik, W. Xu, and C. Theobalt, “Physcap: Physically plausible monocular 3d motion capture in real time,” ACM Transactions on Graphics (ToG) , vol. 39, no. 6, pp. 1–16, 2020
2020
Earlier work this paper cites.
M. Contributors, “Openmmlab’s next generation video understanding toolbox and benchmark,” https://github.com/open-mmlab/mmaction2 , 2020
2020
Earlier work this paper cites.
A. Ghosh, N. Cheema, C. Oguz, C. Theobalt, and P. Slusallek, “Synthesis of compositional animations from textual descriptions,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 1396–1406
2021
Earlier work this paper cites.
A. R. Punnakkal, A. Chandrasekaran, N. Athanasiou, A. Quiros-Ramirez, and M. J. Black, “Babel: Bodies, action and behavior with english labels,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 722–731
2021
Earlier work this paper cites.
K. Xie, T. Wang, U. Iqbal, Y. Guo, S. Fidler, and F. Shkurti, “Physics-based human motion estimation and synthesis from videos,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 11 532–11 541
2021
Earlier work this paper cites.
Z. Luo, R. Hachiuma, Y. Yuan, and K. Kitani, “Dynamics-regulated kinematic policy for egocentric pose estimation,” Advances in Neural Information Processing Systems , vol. 34, pp. 25 019–25 032, 2021
2021
Earlier work this paper cites.
Y. Yuan, S.-E. Wei, T. Simon, K. Kitani, and J. Saragih, “Simpoe: Simulated character control for 3d human pose estimation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 7159–7169
2021
Earlier work this paper cites.
R. Dabral, S. Shimada, A. Jain, C. Theobalt, and V. Golyanik, “Gravity-aware monocular 3d human-object reconstruction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 12 365–12 374
2021
Earlier work this paper cites.
S. Shimada, V. Golyanik, W. Xu, P. Pérez, and C. Theobalt, “Neural monocular 3d human motion capture with physical awareness,” ACM Transactions on Graphics (ToG) , vol. 40, no. 4, pp. 1–15, 2021
2021
Cited alongside, same era.
S. Zhang, Y. Zhang, F. Bogo, M. Pollefeys, and S. Tang, “Learning motion priors for 4d human body capture in 3d scenes,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 11 343–11 353
2021
Cited alongside, same era.
D. Rempe, T. Birdal, A. Hertzmann, J. Yang, S. Sridhar, and L. J. Guibas, “Humor: 3d human motion model for robust pose estimation,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 11 488–11 499
2021
Cited alongside, same era.
H. Zhao, L. Jiang, J. Jia, P. H. Torr, and V. Koltun, “Point transformer,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 16 259–16 268
2021
Cited alongside, same era.
2023
Later among the works it cites.
X. Chen, B. Jiang, W. Liu, Z. Huang, B. Fu, T. Chen, and G. Yu, “Executing your commands via motion diffusion in latent space,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 18 000–18 010
2023
Later among the works it cites.
A. Ghosh, R. Dabral, V. Golyanik, C. Theobalt, and P. Slusallek, “Imos: Intent-driven full-body motion synthesis for human-object interactions,” in Computer Graphics Forum , vol. 42, no. 2. Wiley Online Library, 2023, pp. 1–12
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Cited alongside, same era.
C.-H. P. Huang, H. Yi, M. Höschle, M. Safroshkin, T. Alexiadis, S. Polikovsky, D. Scharstein, and M. J. Black, “Capturing and inferring dense full-body human-scene contact,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 13 274–13 285
2022
Cited alongside, same era.
C. Guo, S. Zou, X. Zuo, S. Wang, W. Ji, X. Li, and L. Cheng, “Generating diverse and natural 3d human motions from text,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 5152–5161
2022
Cited alongside, same era.
G. Tevet, B. Gordon, A. Hertz, A. H. Bermano, and D. Cohen-Or, “Motionclip: Exposing human motion generation to clip space,” in European Conference on Computer Vision . Springer, 2022, pp. 358–374
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Z. Wang, Y. Chen, T. Liu, Y. Zhu, W. Liang, and S. Huang, “Humanise: Language-conditioned human motion generation in 3d scenes,” Advances in Neural Information Processing Systems , vol. 35, pp. 14 959–14 971, 2022
2022
Cited alongside, same era.
M. Petrovich, M. J. Black, and G. Varol, “Temos: Generating diverse human motions from textual descriptions,” in European Conference on Computer Vision . Springer, 2022, pp. 480–497
2022
Cited alongside, same era.
2023
Later among the works it cites.
Z. Zhou and B. Wang, “Ude: A unified driving engine for human motion generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 5632–5641
2023
Later among the works it cites.
R. Dabral, M. H. Mughal, V. Golyanik, and C. Theobalt, “Mofusion: A framework for denoising-diffusion-based motion synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 9760–9770
2023
Later among the works it cites.
J. Kim, J. Kim, and S. Choi, “Flame: Free-form language-based motion synthesis & editing,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 7, 2023, pp. 8255–8263
2023
Later among the works it cites.
S. Ma, Q. Cao, H. Yi, J. Zhang, and D. Tao, “Grammar: Ground-aware motion model for 3d human motion reconstruction,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 2817–2828
2023
Later among the works it cites.
S. Tripathi, A. Chatterjee, J.-C. Passy, H. Yi, D. Tzionas, and M. J. Black, “Deco: Dense estimation of 3d human-scene contact in the wild,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 8001–8013
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Xu, Z. Li, Y.-X. Wang, and L.-Y. Gui, “Interdiff: Generating 3d human-object interactions with physics-informed diffusion,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 14 928–14 940
2023
Later among the works it cites.
J. Li, J. Wu, and C. K. Liu, “Object motion guided human motion synthesis,” ACM Transactions on Graphics (TOG) , vol. 42, no. 6, pp. 1–11, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
C. Diller and A. Dai, “Cg-hoi: Contact-guided 3d human-object interaction generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 19 888–19 901
2024
Closest in time.
Z. Wang, Y. Chen, B. Jia, P. Li, J. Zhang, J. Zhang, T. Liu, Y. Zhu, W. Liang, and S. Huang, “Move as you say interact as you can: Language-guided human motion generation with scene affordance,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 433–444
2024
Closest in time.
C. Guo, X. Zuo, S. Wang, S. Zou, Q. Sun, A. Deng, M. Gong, and L. Cheng, “Action2motion: Conditioned generation of 3d human motions,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 2021–2029
2029
Closest in time.