Fetching the paper…
Reading the bibliography…
We have recently seen tremendous progress in diffusion advances for generating realistic human motions.
Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, et al. (2020) Language models are few-shot learners. Advances in neural information processing systems 33:1877–1901
1901
Earlier work this paper cites.
Bregler C, Malik J (1998) Tracking people with twists and exponential maps. In: Proceedings. 1998 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (Cat. No. 98CB36231), IEEE, pp 8–15
1998
Earlier work this paper cites.
Anguelov D, Srinivasan P, Koller D, Thrun S, Rodgers J, Davis J (2005) Scape: shape completion and animation of people. In: ACM SIGGRAPH 2005 Papers, pp 408–416
2005
Earlier work this paper cites.
Vlasic D, Adelsberger R, Vannucci G, Barnwell J, Gross M, Matusik W, Popović J (2007) Practical motion capture in everyday surroundings. ACM transactions on graphics (TOG) 26(3):35–es
2007
Earlier work this paper cites.
De Aguiar E, Stoll C, Theobalt C, Ahmed N, Seidel HP, Thrun S (2008) Performance capture from sparse multi-view video. In: ACM SIGGRAPH 2008 papers, pp 1–10
2008
Earlier work this paper cites.
Gall J, Rosenhahn B, Brox T, Seidel HP (2010) Optimization and filtering for human motion capture. International Journal of Computer Vision (IJCV) 87(1–2):75–92
2010
Earlier work this paper cites.
Theobalt C, de Aguiar E, Stoll C, Seidel HP, Thrun S (2010) Performance capture from multi-view video. In: Image and Geometry Processing for 3-D Cinematography, Springer, pp 127–149
2010
Earlier work this paper cites.
Van der Aa N, Luo X, Giezeman GJ, Tan RT, Veltkamp RC (2011) Umpm benchmark: A multi-person dataset with synchronized video and motion capture data for evaluation of articulated human motion and interaction. In: 2011 IEEE international conference on computer vision workshops (ICCV Workshops), IEEE, pp 1264–1269
2011
Earlier work this paper cites.
Stoll C, Hasler N, Gall J, Seidel HP, Theobalt C (2011) Fast articulated motion tracking using a sums of Gaussians body model. In: International Conference on Computer Vision (ICCV)
2011
Earlier work this paper cites.
Helten T, Muller M, Seidel HP, Theobalt C (2013) Real-time body tracking with one depth camera and inertial sensors. In: Proceedings of the IEEE international conference on computer vision, pp 1105–1112
2013
Earlier work this paper cites.
Kingma DP, Welling M (2013) Auto-encoding variational bayes. arXiv preprint arXiv:13126114
2013
Earlier work this paper cites.
Liu Y, Gall J, Stoll C, Dai Q, Seidel HP, Theobalt C (2013) Markerless motion capture of multiple characters using multiview image segmentation. IEEE transactions on pattern analysis and machine intelligence 35(11):2720–2735
2013
Earlier work this paper cites.
Loper M, Mahmood N, Romero J, Pons-Moll G, Black MJ (2015) Smpl: A skinned multi-person linear model. ACM transactions on graphics (TOG) 34(6):1–16
2015
Earlier work this paper cites.
Rezende D, Mohamed S (2015) Variational inference with normalizing flows. In: International conference on machine learning, PMLR, pp 1530–1538
2015
Earlier work this paper cites.
Andrews S, Huerta I, Komura T, Sigal L, Mitchell K (2016) Real-time physics-based motion capture with sparse sensors. In: Proceedings of the 13th European conference on visual media production (CVMP 2016), pp 1–10
2016
Earlier work this paper cites.
Bogo F, Kanazawa A, Lassner C, Gehler P, Romero J, Black MJ (2016) Keep it smpl: Automatic estimation of 3d human pose and shape from a single image. In: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14, Springer, pp 561–578
2016
Earlier work this paper cites.
Plappert M, Mandery C, Asfour T (2016) The kit motion-language dataset. Big data 4(4):236–252
2016
Earlier work this paper cites.
Robertini N, Casas D, Rhodin H, Seidel HP, Theobalt C (2016) Model-based outdoor performance capture. In: 2016 Fourth International Conference on 3D Vision (3DV), IEEE, pp 166–175
2016
Earlier work this paper cites.
Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S (2017) Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
Huang Y, Bogo F, Lassner C, Kanazawa A, Gehler PV, Romero J, Akhter I, Black MJ (2017) Towards accurate marker-less human shape and pose estimation over time. In: 2017 international conference on 3D vision (3DV), IEEE, pp 421–430
2017
Earlier work this paper cites.
Lassner C, Romero J, Kiefel M, Bogo F, Black MJ, Gehler PV (2017) Unite the people: Closing the loop between 3d and 2d human representations. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 6050–6059
2017
Earlier work this paper cites.
Malleson C, Gilbert A, Trumble M, Collomosse J, Hilton A, Volino M (2017) Real-time full-body motion capture from video and imus. In: 2017 international conference on 3D vision (3DV), IEEE, pp 449–457
2017
Earlier work this paper cites.
Pavlakos G, Zhou X, Derpanis KG, Daniilidis K (2017) Harvesting multiple views for marker-less 3d human pose annotations. In: Computer Vision and Pattern Recognition (CVPR)
2017
Earlier work this paper cites.
Simon T, Joo H, Matthews I, Sheikh Y (2017) Hand keypoint detection in single images using multiview bootstrapping. In: Computer Vision and Pattern Recognition (CVPR)
2017
Earlier work this paper cites.
Von Marcard T, Rosenhahn B, Black MJ, Pons-Moll G (2017) Sparse inertial poser: Automatic 3d human pose estimation from sparse imus. In: Computer Graphics Forum, Wiley Online Library, vol 36, pp 349–360
2017
Earlier work this paper cites.
Ahn H, Ha T, Choi Y, Yoo H, Oh S (2018) Text2action: Generative adversarial synthesis from language to action. In: 2018 IEEE International Conference on Robotics and Automation (ICRA), IEEE, pp 5915–5920
2018
Earlier work this paper cites.
Huang Y, Kaufmann M, Aksan E, Black MJ, Hilliges O, Pons-Moll G (2018) Deep inertial poser: Learning to reconstruct human pose from sparse inertial measurements in real time. ACM Transactions on Graphics (TOG) 37(6):1–15
2018
Earlier work this paper cites.
Loshchilov I, Hutter F (2018) Decoupled weight decay regularization. In: International Conference on Learning Representations
2018
Earlier work this paper cites.
Von Marcard T, Henschel R, Black MJ, Rosenhahn B, Pons-Moll G (2018) Recovering accurate 3d human pose in the wild using imus and a moving camera. In: Proceedings of the European conference on computer vision (ECCV), pp 601–617
2018
Earlier work this paper cites.
Zheng Z, Yu T, Li H, Guo K, Dai Q, Fang L, Liu Y (2018) Hybridfusion: Real-time performance capture using a single depth sensor and sparse imus. In: Proceedings of the European Conference on Computer Vision (ECCV), pp 384–400
2018
Earlier work this paper cites.
Ahuja C, Morency LP (2019) Language2pose: Natural language grounded pose forecasting. In: 2019 International Conference on 3D Vision (3DV), IEEE, pp 719–728
2019
Earlier work this paper cites.
Gilbert A, Trumble M, Malleson C, Hilton A, Collomosse J (2019) Fusing visual and inertial sensors with semantics for 3d human pose estimation. International Journal of Computer Vision 127:381–397
2019
Earlier work this paper cites.
Habermann M, Xu W, Zollhöfer M, Pons-Moll G, Theobalt C (2019) Livecap: Real-time human performance capture from monocular video. ACM Transactions on Graphics (TOG) 38(2):14:1–14:17
2019
Cited alongside, same era.
Kanazawa A, Zhang JY, Felsen P, Malik J (2019) Learning 3d human dynamics from video. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 5614–5623
2019
Cited alongside, same era.
Kenton JDMWC, Toutanova LK (2019) Bert: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of naacL-HLT, vol 1, p 2
2019
Cited alongside, same era.
Kolotouros N, Pavlakos G, Black MJ, Daniilidis K (2019) Learning to reconstruct 3d human pose and shape via model-fitting in the loop. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 2252–2261
2019
Cited alongside, same era.
Wang J, Yan S, Dai B, Lin D (2021) Scene-aware generative network for human motion synthesis. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 12206–12215
2021
Later among the works it cites.
Yi X, Zhou Y, Xu F (2021) Transpose: Real-time 3d human translation and pose estimation with six inertial sensors. ACM Transactions on Graphics (TOG) 40(4):1–13
2021
Later among the works it cites.
Zanfir A, Bazavan EG, Zanfir M, Freeman WT, Sukthankar R, Sminchisescu C (2021) Neural descent for visual 3d human pose and shape. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 14484–14493
2021
Later among the works it cites.
Ao T, Gao Q, Lou Y, Chen B, Liu L (2022) Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings. ACM Transactions on Graphics (TOG) 41(6):1–19
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lee HY, Yang X, Liu MY, Wang TC, Lu YD, Yang MH, Kautz J (2019) Dancing to music. Advances in neural information processing systems 32
2019
Cited alongside, same era.
Liu J, Shahroudy A, Perez M, Wang G, Duan LY, Kot AC (2019) Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding. IEEE transactions on pattern analysis and machine intelligence 42(10):2684–2701
2019
Cited alongside, same era.
Malleson C, Collomosse J, Hilton A (2019) Real-time multi-person motion capture from multi-view video and imus. International Journal of Computer Vision pp 1–18
2019
Cited alongside, same era.
Pavlakos G, Choutas V, Ghorbani N, Bolkart T, Osman AA, Tzionas D, Black MJ (2019) Expressive body capture: 3d hands, face, and body from a single image. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10975–10985
2019
Cited alongside, same era.
Starke S, Zhang H, Komura T, Saito J (2019) Neural state machine for character-scene interactions. ACM Trans Graph 38(6):209–1
2019
Cited alongside, same era.
Vicon (2019) Vicon Motion Systems
2019
Cited alongside, same era.
Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2020) Generative adversarial networks. Communications of the ACM 63(11):139–144
2020
Cited alongside, same era.
Habermann M, Xu W, Zollhofer M, Pons-Moll G, Theobalt C (2020) Deepcap: Monocular human performance capture using weak supervision. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2020
Cited alongside, same era.
Athanasiou N, Petrovich M, Black MJ, Varol G (2022) Teach: Temporal action composition for 3d humans. In: 2022 International Conference on 3D Vision (3DV), IEEE, pp 414–423
2022
Later among the works it cites.
Chen X, Su Z, Yang L, Cheng P, Xu L, Fu B, Yu G (2022) Learning variational motion prior for video-based motion capture. arXiv preprint arXiv:221015134
2022
Later among the works it cites.
Guo C, Zuo X, Wang S, Cheng L (2022b) Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts. In: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXV, Springer, pp 580–597
2022
Later among the works it cites.
Habibie I, Elgharib M, Sarkar K, Abdullah A, Nyatsanga S, Neff M, Theobalt C (2022) A motion matching-based framework for controllable gesture synthesis from speech. In: ACM SIGGRAPH 2022 Conference Proceedings, pp 1–9
2022
Later among the works it cites.
Li B, Zhao Y, Zhelun S, Sheng L (2022) Danceformer: Music conditioned 3d dance generation with parametric motion transformer. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 36, pp 1272–1279
2022
Later among the works it cites.
Lucas T, Baradel F, Weinzaepfel P, Rogez G (2022) Posegpt: Quantization-based 3d human motion generation and forecasting. In: European Conference on Computer Vision, Springer, pp 417–435
2022
Later among the works it cites.
Petrovich M, Black MJ, Varol G (2022) Temos: Generating diverse human motions from textual descriptions. In: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXII, Springer, pp 480–497
2022
Later among the works it cites.
Starke S, Mason I, Komura T (2022) Deepphase: Periodic autoencoders for learning motion phase manifolds. ACM Transactions on Graphics (TOG) 41(4):1–13
2022
Later among the works it cites.
Tevet G, Gordon B, Hertz A, Bermano AH, Cohen-Or D (2022a) Motionclip: Exposing human motion generation to clip space. In: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXII, Springer, pp 358–374
2022
Later among the works it cites.
Wang Z, Chen Y, Liu T, Zhu Y, Liang W, Huang S (2022) Humanise: Language-conditioned human motion generation in 3d scenes. Advances in Neural Information Processing Systems 35:14959–14971
2022
Later among the works it cites.
Yi X, Zhou Y, Habermann M, Shimada S, Golyanik V, Theobalt C, Xu F (2022) Physical inertial poser (pip): Physics-aware real-time human motion tracking from sparse inertial sensors. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
Later among the works it cites.
Zhang M, Cai Z, Pan L, Hong F, Guo X, Yang L, Liu Z (2022) Motiondiffuse: Text-driven human motion generation with diffusion model. arXiv preprint arXiv:220815001
2022
Later among the works it cites.
Chen X, Jiang B, Liu W, Huang Z, Fu B, Chen T, Yu G (2023) Executing your commands via motion diffusion in latent space. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 18000–18010
2023
Closest in time.
Jiang B, Chen X, Liu W, Yu J, Yu G, Chen T (2023) Motiongpt: Human motion as a foreign language. arXiv preprint arXiv:230614795
2023
Closest in time.
Kalakonda SS, Maheshwari S, Sarvadevabhatla RK (2023) Action-gpt: Leveraging large-scale language models for improved and generalized action generation. In: 2023 IEEE International Conference on Multimedia and Expo (ICME), IEEE, pp 31–36
2023
Closest in time.
Kim J, Kim J, Choi S (2023) Flame: Free-form language-based motion synthesis & editing. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 37, pp 8255–8263
2023
Closest in time.
Liang H, He Y, Zhao C, Li M, Wang J, Yu J, Xu L (2023) Hybridcap: Inertia-aid monocular capture of challenging human motions. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 37, pp 1539–1548
2023
Closest in time.
Movella (2022) Movella xsens products
2023
Closest in time.
OpenAI (2023) Gpt-4 technical report
2023
Closest in time.
Ren Y, Zhao C, He Y, Cong P, Liang H, Yu J, Xu L, Ma Y (2023) Lidar-aid inertial poser: Large-scale human motion capture by sparse inertial and lidar sensors. IEEE Transactions on Visualization and Computer Graphics 29(5):2337–2347
2023
Closest in time.
Shafir Y, Tevet G, Kapon R, Bermano AH (2023) Human motion diffusion as a generative prior. arXiv preprint arXiv:230301418
2023
Closest in time.
Tanaka M, Fujiwara K (2023) Role-aware interaction generation from textual description. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 15999–16009
2023
Closest in time.
Touvron H, Lavril T, Izacard G, Martinet X, Lachaux MA, Lacroix T, Rozière B, Goyal N, Hambro E, Azhar F, et al. (2023) Llama: Open and efficient foundation language models. arXiv preprint arXiv:230213971
2023
Closest in time.
Xu L, Song Z, Wang D, Su J, Fang Z, Ding C, Gan W, Yan Y, Jin X, Yang X, et al. (2023) Actformer: A gan-based transformer towards general action-conditioned 3d human motion generation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 2228–2238
2023
Closest in time.
Yuan Y, Song J, Iqbal U, Vahdat A, Kautz J (2023) Physdiff: Physics-guided human motion diffusion model. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 16010–16021
2023
Closest in time.
Z-cam (2022) Z CAM Cinema Camera
2023
Closest in time.
Guo C, Zuo X, Wang S, Zou S, Sun Q, Deng A, Gong M, Cheng L (2020) Action2motion: Conditioned generation of 3d human motions. In: Proceedings of the 28th ACM International Conference on Multimedia, pp 2021–2029
2029
Closest in time.