Fetching the paper…
Reading the bibliography…
Vision-language-action models (VLAs) have shown generalization capabilities in robotic manipulation tasks by inheriting from vision-language models (VLMs) and learning action generation.
I. A. Şucan, M. Moll, and L. E. Kavraki, “The Open Motion Planning Library,” IEEE Robotics & Automation Magazine , vol. 19, no. 4, pp. 72–82, December 2012, https://ompl.kavrakilab.org
2012
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 652–660
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison, “Rlbench: The robot learning benchmark & learning environment,” IEEE Robotics and Automation Letters , vol. 5, no. 2, pp. 3019–3026, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn, “Robonet: Large-scale multi-robot learning,” in Conference on Robot Learning . PMLR, 2020, pp. 885–897
2020
Earlier work this paper cites.
2022
Earlier work this paper cites.
Y. Zhu, P. Stone, and Y. Zhu, “Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation,” IEEE Robotics and Automation Letters , vol. 7, no. 2, pp. 4126–4133, 2022
2022
Earlier work this paper cites.
S. Nasiriany, T. Gao, A. Mandlekar, and Y. Zhu, “Learning and retrieval from prior data for skill-based imitation learning,” in Conference on Robot Learning (CoRL) , 2022
2022
Earlier work this paper cites.
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn, “Bc-z: Zero-shot task generalization with robotic imitation learning,” in Conference on Robot Learning . PMLR, 2022, pp. 991–1002
2022
Earlier work this paper cites.
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine, “Bridge data: Boosting generalization of robotic skills with cross-domain datasets,” in RSS , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
E. Rosete-Beas, O. Mees, G. Kalweit, J. Boedecker, and W. Burgard, “Latent plans for task agnostic offline reinforcement learning,” in Proceedings of the 6th Conference on Robot Learning (CoRL) , 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu et al. , “Palm-e: an embodied multimodal language model,” in Proceedings of the 40th International Conference on Machine Learning , 2023, pp. 8469–8488
2023
Earlier work this paper cites.
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” in Proceedings of Robotics: Science and Systems (RSS) , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in International Conference on Computer Vision (ICCV) , 2023
2023
Earlier work this paper cites.
H. Liu, S. Nasiriany, L. Zhang, Z. Bao, and Y. Zhu, “Robot learning on the job: Human-in-the-loop autonomy and learning during deployment,” in Robotics: Science and Systems (RSS) , 2023
2023
Earlier work this paper cites.
H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V. Myers, K. Fang, C. Finn, and S. Levine, “Bridgedata v2: A dataset for robot learning at scale,” 2023
2023
Earlier work this paper cites.
R. Mendonca, S. Bahl, and D. Pathak, “Structured world models from human videos,” CoRL , 2023
2023
Cited alongside, same era.
A. Padalkar, G. Quere, A. Raffin, J. Silvério, and F. Stulp, “A guided reinforcement learning approach using shared control templates for learning manipulation skills in the real world,” 2023
2023
Cited alongside, same era.
N. M. M. Shafiullah, A. Rai, H. Etukuru, Y. Liu, I. Misra, S. Chintala, and L. Pinto, “On bringing robots home,” 2023
2023
Cited alongside, same era.
J. Luo, C. Xu, F. Liu, L. Tan, Z. Lin, J. Wu, P. Abbeel, and S. Levine, “FMB: A functional manipulation benchmark for generalizable robotic learning,” https://functional-manipulation-benchmark.github.io , 2023
2023
Cited alongside, same era.
M. Heo, Y. Lee, D. Lee, and J. J. Lim, “Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation,” in Robotics: Science and Systems , 2023
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
S. Dass, J. Yapeter, J. Zhang, J. Zhang, K. Pertsch, S. Nikolaidis, and J. J. Lim, “CLVR jaco play dataset,” 2023. [Online]. Available: https://github.com/clvrai/clvr_jaco_play_dataset
2023
Cited alongside, same era.
C. Lynch, A. Wahid, J. Tompson, T. Ding, J. Betker, R. Baruch, T. Armstrong, and P. Florence, “Interactive language: Talking to robots in real time,” IEEE Robotics and Automation Letters , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
V. Kumar, R. Shah, G. Zhou, V. Moens, V. Caggiano, A. Gupta, and A. Rajeswaran, “Robohive: A unified framework for robot learning,” in Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track , 2023
2023
Cited alongside, same era.
S. Belkhale, Y. Cui, and D. Sadigh, “Hydra: Hybrid robot actions for imitation learning,” arxiv , 2023
2023
Cited alongside, same era.
O. Mees, J. Borja-Diaz, and W. Burgard, “Grounding language with visual affordances over unstructured data,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) , London, UK, 2023
2023
Cited alongside, same era.
G. Zhou, V. Dean, M. K. Srirama, A. Rajeswaran, J. Pari, K. Hatch, A. Jain, T. Yu, P. Abbeel, L. Pinto, C. Finn, and A. Gupta, “Train offline, test online: A real robot learning benchmark,” 2023
2023
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn et al. , “Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 1702–1713
2025
Closest in time.
K. Wu, C. Hou, J. Liu, Z. Che, X. Ju et al. , “Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation,” in Robotics: Science and Systems (RSS) 2025 . Robotics: Science and Systems Foundation, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Jia, J. Liu, S. Chen, C. Gu, Z. Wang, L. Luo, X. Li, P. Wang, Z. Wang, R. Zhang et al. , “Lift3d policy: Lifting 2d foundation models for robust 3d robotic manipulation,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 17 347–17 358
2025
Closest in time.
Y. Tang, R. Zhang, J. Liu, Z. Guo, B. Zhao, Z. Wang, P. Gao, H. Li, D. Wang, and X. Li, “Any2point: Empowering any-modality large models for efficient 3d understanding,” in European Conference on Computer Vision . Springer, 2025, pp. 456–473
2025
Closest in time.
2025
Closest in time.