Fetching the paper…
Reading the bibliography…
This paper addresses the limitations of current humanoid robot control frameworks, which primarily rely on reactive mechanisms and lack autonomous interaction capabilities due to data scarcity.
Latent structured models for human pose estimation
Catalin Ionescu, Fuxin Li, C. S · 2011
Earlier work this paper cites.
Whole-body motion planning with centroidal dynamics and full kinematics
Dai, H., Valenzuela, A., and Tedrake, R · 2014
Earlier work this paper cites.
Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Ionescu, C., Papava, D., Olaru, V., and Sminchisescu, C · 2014
Earlier work this paper cites.
Optimization-based locomotion planning, estimation, and control design for the atlas humanoid robot
Kuindersma, S., Deits, R., Fallon, M., Valenzuela, A., Dai, H., Permenter, F., Koolen, T., Marion, P., and Tedrake, R · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Amass: Archive of motion capture as surface shapes
Mahmood, N., Ghorbani, N., Troje, N. F., Pons-Moll, G., and Black, M. J · 2019
Earlier work this paper cites.
Action2motion: Conditioned generation of 3d human motions
Guo, C., Zuo, X., Wang, S., Zou, S., Sun, Q., Deng, A., Gong, M., and Cheng, L · 2020
Earlier work this paper cites.
Whole body model predictive control with a memory of motion: Experiments on a torque-controlled talos
Dantec, E., Budhiraja, R., Roig, A., Lembono, T., Saurel, G., Stasse, O., Fernbach, P., Tonneau, S., Vijayakumar, S., Calinon, S., et al · 2021
Earlier work this paper cites.
Isaac gym: High performance gpu-based physics simulation for robot learning
Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., et al · 2021
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Cited alongside, same era.
Online non-linear centroidal mpc for humanoid robots payload carrying with contact-stable force parametrization
Elobaid, M., Romualdi, G., Nava, G., Rapetti, L., Mohamed, H. A. O., and Pucci, D · 2023
Cited alongside, same era.
Dynamic loco-manipulation on hector: Humanoid for enhanced control and open-source research
Li, J., Ma, J., Kolt, O., Shah, M., and Nguyen, Q · 2023
Cited alongside, same era.
Motion-x: A large-scale 3d expressive whole-body human motion dataset
Lin, J., Zeng, A., Lu, S., Cai, Y., Zhang, R., Wang, H., and Zhang, L · 2023
Cited alongside, same era.
Visual instruction tuning, 2023
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Exbody2: Advanced expressive humanoid whole-body control
Ji, M., Peng, X., Liu, F., Li, J., Yang, G., Cheng, X., and Wang, X · 2024
Later among the works it cites.
Harmon: Whole-body motion generation of humanoid robots from language descriptions
Jiang, Z., Xie, Y., Li, J., Yuan, Y., Zhu, Y., and Zhu, Y · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
Kim, M. J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., et al · 2024
Later among the works it cites.
Vila: On pre-training for visual language models
Lin, J., Yin, H., Ping, W., Molchanov, P., Shoeybi, M., and Han, S · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Smpl: A skinned multi-person linear model
Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., and Black, M. J · 2023
Cited alongside, same era.
Perpetual humanoid control for real-time simulated avatars
Luo, Z., Cao, J., Kitani, K., Xu, W., et al · 2023
Cited alongside, same era.
Human motion diffusion model
Tevet, G., Raab, S., Gordon, B., Shafir, Y., Cohen-or, D., and Bermano, A. H · 2023
Cited alongside, same era.
p i _ 0 pi\_0 : A vision-language-action flow model for general robot control
Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., et al · 2024
Cited alongside, same era.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
Cheang, C.-L., Chen, G., Jing, Y., Kong, T., Li, H., Li, Y., Liu, Y., Wu, H., Xu, J., Yang, Y., et al · 2024
Cited alongside, same era.
Expressive whole-body control for humanoid robots
Cheng, X., Ji, Y., Chen, J., Yang, R., Yang, G., and Wang, X · 2024
Cited alongside, same era.
Generating diverse and natural 3d human motions from text
Guo, C., Zou, S., Zuo, X., Wang, S., Ji, W., Li, X., and Cheng, L
Cited in the paper.
Later among the works it cites.
Mobile-television: Predictive motion priors for humanoid whole-body control
Lu, C., Cheng, X., Li, J., Yang, S., Ji, M., Yuan, C., Yang, G., Yi, S., and Wang, X · 2024
Later among the works it cites.
Learning from massive human videos for universal humanoid pose control
Mao, J., Zhao, S., Song, S., Shi, T., Ye, J., Zhang, M., Geng, H., Malik, J., Guizilini, V., and Wang, Y · 2024
Later among the works it cites.
Quart-online: Latency-free large multimodal language model for quadruped robot learning
Tong, X., Ding, P., Wang, D., Zhang, W., Cui, C., Sun, M., Fan, Y., Zhao, H., Zhang, H., Dang, Y., et al · 2024
Later among the works it cites.
Humanvla: Towards vision-language directed object rearrangement by physical humanoid
Xu, X., Zhang, Y., Li, Y.-L., Han, L., and Lu, C · 2024
Later among the works it cites.
Quar-vla: Vision-language-action model for quadruped robots
Ding, P., Zhao, H., Zhang, W., Song, W., Zhang, M., Huang, S., Yang, N., and Wang, D · 2025
Closest in time.
Tram: Global trajectory and motion of 3d humans from in-the-wild videos
Wang, Y., Wang, Z., Liu, L., and Daniilidis, K · 2025
Closest in time.