Fetching the paper…
Reading the bibliography…
Building autonomous robotic agents capable of achieving human-level performance in real-world embodied tasks is an ultimate goal in humanoid robot research.
Dynamic walk of a biped
Miura, H. and Shimoyama, I · 1984
Earlier work this paper cites.
Whole body humanoid control from human motion descriptors
Dariush, B., Gienger, M., Jian, B., Goerick, C., and Fujimura, K · 2008
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., Bernstein, M. S., and Li, F.-F · 2016
Earlier work this paper cites.
Humanoid robotics: a reference
Goswami, A. and Vadakkepat, P · 2018
Earlier work this paper cites.
Learning locomotion skills for cassie: Iterative design and sim-to-real
Xie, Z., Clary, P., Dao, J., Morais, P., Hurst, J., and Panne, M · 2020
Earlier work this paper cites.
Isaac gym: High performance gpu-based physics simulation for robot learning
Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., et al · 2021
Earlier work this paper cites.
Blind bipedal stair traversal via sim-to-real reinforcement learning
Siekmann, J., Green, K., Warila, J., Fern, A., and Hurst, J · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models
Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., et al · 2022
Earlier work this paper cites.
Vima: General robot manipulation with multimodal prompts
Jiang, Y., Gupta, A., Zhang, Z., Wang, G., Dou, Y., Chen, Y., Fei-Fei, L., Anandkumar, A., Zhu, Y., and Fan, L · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Earlier work this paper cites.
Minigpt-v2: large language model as a unified interface for vision-language multi-task learning
Chen, J., Zhu, D., Shen, X., Li, X., Liu, Z., Zhang, P., Krishnamoorthi, R., Chandra, V., Xiong, Y., and Elhoseiny, M · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Xu, Z., Feng, S., Cousineau, E., Du, Y., Burchfiel, B., Tedrake, R., and Song, S · 2023
Earlier work this paper cites.
Du, Y., Yang, M., Florence, P., Xia, F., Wahid, A., Ichter, B., Sermanet, P., Yu, T., Abbeel, P., Tenenbaum, J. B., et al · 2023
Cited alongside, same era.
Foundation models in robotics: Applications, challenges, and the future
Firoozi, R., Tucker, J., Tian, S., Majumdar, A., Sun, J., Liu, W., Zhu, Y., Song, S., Kapoor, A., Hausman, K., et al · 2023
Cited alongside, same era.
Toward general-purpose robots via foundation models: A survey and meta-analysis
Hu, Y., Xie, Q., Jain, V., Francis, J., Patrikar, J., Keetha, N., Kim, S., Xie, Y., Zhang, T., Fang, H.-S., et al · 2023
Cited alongside, same era.
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S. C. H · 2023
Cited alongside, same era.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Efficient residual learning with mixture-of-experts for universal dexterous grasping
Huang, Z., Yuan, H., Fu, Y., and Lu, Z · 2024
Later among the works it cites.
Exbody2: Advanced expressive humanoid whole-body control
Ji, M., Peng, X., Liu, F., Li, J., Yang, G., Cheng, X., and Wang, X · 2024
Later among the works it cites.
Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization
Jin, Y., Sun, Z., Xu, K., Xu, K., Chen, L., Jiang, H., Huang, Q., Song, C., Liu, Y., Zhang, D., Song, Y., Gai, K., and Mu, Y · 2024
Later among the works it cites.
Smart-llm: Smart multi-agent robot task planning using large language models
Kannan, S. S., Venkatesh, V. L., and Min, B.-C · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Cited alongside, same era.
Audio-visual llm for video understanding
Shu, F., Zhang, L., Jiang, H., and Xie, C · 2023
Cited alongside, same era.
Optimization-based control for dynamic legged robots
Wensing, P. M., Posa, M., Hu, Y., Escande, A., Mansard, N., and Del Prete, A · 2023
Cited alongside, same era.
Learning interactive real-world simulators
Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., and Abbeel, P · 2023
Cited alongside, same era.
Skill reinforcement learning and planning for open-world long-horizon tasks
Yuan, H., Zhang, C., Wang, H., Xie, F., Cai, P., Dong, H., and Lu, Z · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L · 2023
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware
Zhao, T. Z., Kumar, V., Levine, S., and Finn, C · 2023
Cited alongside, same era.
Kim, M. J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., et al · 2024
Later among the works it cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H. A., Khan, S., and Khan, F · 2024
Later among the works it cites.
Real-world humanoid locomotion with reinforcement learning
Radosavovic, I., Xiao, T., Zhang, B., Darrell, T., Malik, J., and Sreenath, K · 2024
Later among the works it cites.
Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning
Shao, H., Qian, S., Xiao, H., Song, G., Zong, Z., Wang, L., Liu, Y., and Li, H · 2024
Later among the works it cites.
Cradle: Empowering foundation agents towards general computer control
Tan, W., Zhang, W., Xu, X., Xia, H., Ding, G., Li, B., Zhou, B., Yue, J., Jiang, J., Li, Y., et al · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
Team, O. M., Ghosh, D., Walke, H., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Kreiman, T., Xu, C., et al · 2024
Later among the works it cites.
Chatterbox: Multi-round multimodal referring and grounding
Tian, Y., Ma, T., Xie, L., Qiu, J., Tang, X., Zhang, Y., Jiao, J., Tian, Q., and Ye, Q · 2024
Later among the works it cites.
A survey on large language model based autonomous agents
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., et al · 2024
Later among the works it cites.
Ace: A cross-platform visual-exoskeletons system for low-cost dexterous teleoperation
Yang, S., Liu, M., Qin, Y., Ding, R., Li, J., Cheng, X., Yang, R., Yi, S., and Wang, X · 2024
Later among the works it cites.
Cross-embodiment dexterous grasping with reinforcement learning
Yuan, H., Zhou, B., Fu, Y., and Lu, Z · 2024
Later among the works it cites.
Learning diverse bimanual dexterous manipulation skills from human demonstrations
Zhou, B., Yuan, H., Fu, Y., and Lu, Z · 2024
Later among the works it cites.
Zhuang, Z., Yao, S., and Zhao, H · 2024
Later among the works it cites.
Gu, Z., Li, J., Shen, W., Yu, W., Xie, Z., McCrory, S., Cheng, X., Shamsah, A., Griffin, R., Liu, C. K., et al · 2025
Closest in time.