Fetching the paper…
Reading the bibliography…
Recent advances in control robot methods, from end-to-end vision-language-action frameworks to modular systems with predefined primitives, have advanced robots' ability to follow natural language instructions.
L. P. Kaelbling, “The foundation of efficient robot learning,” Science , vol. 369, no. 6506, pp. 915–916, 2020
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PmLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
D. Shah, B. Osiński, B. Ichter, Y. Hu, A. Saddiqi et al. , “Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,” in Conference on Robot Learning (CoRL) , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu et al. , “Rt-1: Robotics transformer for real-world control at scale,” arXiv , 2022
2022
Earlier work this paper cites.
W. Huang, F. Xia, J. Tompson, A. Zeng, B. Ichter et al. , “Inner monologue: Embodied reasoning through planning with language models,” in Conference on Robot Learning (CoRL) , 2022
2022
Earlier work this paper cites.
S. Nair, S. Khandelwal, M. Danielczuk, K. Xu, P. Florence et al. , “R3m: A universal visual representation for robot manipulation,” in Conference on Robot Learning (CoRL) , 2022
2022
Earlier work this paper cites.
M. Ahn, A. Brohan, N. Brown, and et al., “Do as i can, not as i say: Grounding language in robotic affordances,” 2022
2022
Earlier work this paper cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International conference on machine learning . PMLR, 2023, pp. 19 730–19 742
2023
Earlier work this paper cites.
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng, “Code as policies: Language model programs for embodied control,” ICRA , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Brohan, N. Brown, J. Carbajal, and et al., “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” 2023
2023
Earlier work this paper cites.
A. Padalkar, A. Pooley, A. Jain, K. Hausman, A. Rai et al. , “Open x-embodiment: Robotic learning datasets and rt-x models,” in Conference on Robot Learning (CoRL) , 2023
2023
Earlier work this paper cites.
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K. Hausman et al. , “Open-world object manipulation using pre-trained vision-language models,” in Conference on Robot Learning (CoRL) , 2023
2023
Cited alongside, same era.
T. Xiao, H. Chan, P. Sermanet, A. Wahid, S. Levine et al. , “Robotic skill acquisition via instruction augmentation with vision-language models,” in Robotics: Science and Systems (RSS) , 2023
2023
Cited alongside, same era.
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, L. Fei-Fei et al. , “Vima: General robot manipulation with multimodal prompts,” in International Conference on Machine Learning (ICML) , 2023
2023
Cited alongside, same era.
W. Huang, F. Xia, D. Shah, A. Zeng, P. Florence et al. , “Grounded decoding: Guiding text generation with grounded models for robot control,” in Advances in Neural Information Processing Systems (NeurIPS) , 2023
2023
Cited alongside, same era.
2024
Later among the works it cites.
N. Ravi, V. Gabeur, Y.-T. Hu, and et al., “Sam 2: Segment anything in images and videos,” 2024
2024
Later among the works it cites.
Z. Yan, S. Li, Z. Wang, L. Wu, H. Wang, J. Zhu, L. Chen, and J. Liu, “Dynamic open-vocabulary 3d scene graphs for long-term language-guided mobile manipulation,” IEEE Robotics and Automation Letters , 2025
2025
Closest in time.
E. Zhou and Q. e. a. Su, “Code-as-monitor: Constraint-aware visual programming for reactive and proactive robotic failure detection,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 6919–6929
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter et al. , “Palm-e: An embodied multimodal language model,” in International Conference on Machine Learning (ICML) , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Cited alongside, same era.
Y. Mu, J. Chen, Q. Zhang, S. Chen, Q. Yu, C. Ge, R. Chen, Z. Liang, M. Hu, C. Tao et al. , “Robocodex: Multimodal code generation for robotic behavior synthesis,” arXiv , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
P. Ding, H. Zhao, W. Zhang, W. Song, S. Huang et al. , “Quar-vla: Vision-language-action model for quadruped robots,” in European Conference on Computer Vision (ECCV) , 2024, pp. 352–367
2024
Cited alongside, same era.
J. Liu, P. An, Z. Liu, R. Zhang, C. Gu et al. , “Robomamba: Efficient vision-language-action model for robotic reasoning and manipulation,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 37, pp. 40 085–40 110, 2024
2024
Cited alongside, same era.
J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 5294–5306
2025
Closest in time.
2025
Closest in time.
R. Mon-Williams, G. Li, R. Long, W. Du, and C. G. Lucas, “Embodied large language models enable robots to complete complex tasks in unpredictable environments,” Nature Machine Intelligence , vol. 7, no. 7, pp. 592–601, 2025
2025
Closest in time.
J. Yang, R. Tan, Q. Wu, R. Zheng, Y. Liang et al. , “Magma: A foundation model for multimodal ai agents,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2025
2025
Closest in time.
G. Kang, J. Kim, K. Shim, J. Lee, and B.-T. Zhang, “Clip-rt: Learning language-conditioned robotic policies from natural language supervision,” in Robotics: Science and Systems (RSS) , 2025
2025
Closest in time.
J. Wen, Y. Zhu, J. Li, M. Zhu, Z. Tang, K. Wu, Z. Xu, N. Liu, R. Cheng, C. Shen et al. , “Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation,” IEEE Robotics and Automation Letters , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Anthropic, “Claude sonnet 4 technical documentation,” https://www.anthropic.com , 2025
2025
Closest in time.