Fetching the paper…
Reading the bibliography…
While Vision-Language-Action (VLA) models have demonstrated impressive capabilities in robotic manipulation, their performance in complex reasoning and long-horizon task planning is limited by data scarcity and model capacity.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, et al. , “Rt-1: Robotics Transformer for Real-World Control at Scale.” in Robotics: Science and Systems Conference (RSS) , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, et al. , “Voxposer: Composable 3d Value Maps for Robotic Manipulation with Language Models.” in Conference on Robot Learning (CoRL) , 2023, pp. 540–562
2023
Earlier work this paper cites.
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, et al. , “Progprompt: Generating Situated Robot Task Plans using Large Language Models.” in IEEE International Conference on Robotics and Automation (ICRA) , vol. 47, no. 8, 2023, pp. 11 523–11 530
2023
Earlier work this paper cites.
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, et al. , “Code as Policies: Language Model Programs for Embodied Control.” in IEEE International Conference on Robotics and Automation (ICRA) , 2023, pp. 9493–9500
2023
Earlier work this paper cites.
K. Rana, J. Haviland, S. Garg, J. Abou-Chakra, I. D. Reid, et al. , “Sayplan: Grounding Large Language Models using 3d Scene Graphs for Scalable Robot Task Planning.” in Conference on Robot Learning (CoRL) , 2023, pp. 23–72
2023
Earlier work this paper cites.
H. Fang, C. Wang, H. Fang, M. Gou, J. Liu, et al. , “Anygrasp: Robust and Efficient Grasp Perception in Spatial and Temporal Domains.” IEEE Transactions on Robotics , vol. 39, no. 5, pp. 3929–3945, 2023
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, et al. , “Toolllm: Facilitating Large Language Models to Master 16000+ Real-world APIs,” in The Twelfth International Conference on Learning Representations , 2024
2024
Cited alongside, same era.
B. Xiao, H. Wu, W. Xu, X. Dai, H. Hu, et al. , “Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024
2024
Later among the works it cites.
M. Ahn, D. Dwibedi, C. Finn, M. G. Arenas, K. Gopalakrishnan, et al. , “Autort: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents,” in First Workshop on Vision-Language Models for Navigation and Manipulation at ICRA 2024 , 2024
2024
Later among the works it cites.
K. Bousmalis, G. Vezzani, D. Rao, C. M. Devin, A. X. Lee, et al. , “Robocat: A Self-Improving Generalist Agent for Robotic Manipulation.” Transactions on Machine Learning Research (TMLR) , vol. 2024, 2024
2024
Later among the works it cites.
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Guo, S. Cheng, H. Wang, S. Liang, Y. Qin, et al. , “Stabletoolbench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models.” in Annual Meeting of the Association for Computational Linguistics (ACL) , 2024, pp. 11 143–11 156
2024
Cited alongside, same era.
2024
Cited alongside, same era.
W. Huang, C. Wang, Y. Li, R. Zhang, and L. Fei-Fei, “Rekep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation,” in 8th Annual Conference on Robot Learning , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
A. Mei, G.-N. Zhu, H. Zhang, and Z. Gan, “Replanvlm: Replanning Robotic Tasks With Visual Language Models.” IEEE Robotics and Automation Letters , vol. 9, no. 11, pp. 10 201–10 208, 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. Lou, K. Xu, Z. Zhou, and R. Xiong, “Explorevlm: Closed-Loop Robot Exploration Task Planning with Vision-Language Models,” arXiv , 2025
2025
Closest in time.