Fetching the paper…
Reading the bibliography…
A fundamental requirement for real-world robotic deployment is the ability to understand and respond to natural language instructions.
E. Coumans and Y. Bai, “Pybullet, a python module for physics simulation for games, robotics and machine learning,” http://pybullet.org , 2016–2019
2019
Earlier work this paper cites.
A. Zacharaki, I. Kostavelis, A. Gasteratos, and I. Dokas, “Safety bounds in human robot interaction: A survey,” Safety science , vol. 127, p. 104667, 2020
2020
Earlier work this paper cites.
Y. Han, G. Huang, S. Song, L. Yang, H. Wang, and Y. Wang, “Dynamic neural networks: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, pp. 7436–7456, 2021
2021
Earlier work this paper cites.
C. Lynch and P. Sermanet, “Language conditioned imitation learning over unstructured data,” Robotics: Science and Systems , 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” IEEE Robotics and Automation Letters , vol. 7, no. 3, pp. 7327–7334, 2022
2022
Earlier work this paper cites.
O. Mees, L. Hermann, and W. Burgard, “What matters in language conditioned robotic imitation learning over unstructured data,” IEEE Robotics and Automation Letters , vol. 7, no. 4, pp. 11 205–11 212, 2022
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations (ICLR) , 2022
2022
Earlier work this paper cites.
A. B. et al., “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” Proceedings of The 7th Conference on Robot Learning (CoRL) , 2023
2023
Earlier work this paper cites.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu et al. , “Rt-1: Robotics transformer for real-world control at scale,” Robotics: Science and Systems , 2023
2023
Earlier work this paper cites.
Z. Xia, D. Han, Y. Han, X. Pan, S. Song, and G. Huang, “Gsva: Generalized segmentation via multimodal large language models,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 3858–3869, 2023
2023
Earlier work this paper cites.
H. Cao, G. Chen, Z. Li, Q. Feng, J. Lin, and A. Knoll, “Efficient grasp detection network with gaussian-based grasp representation for robotic manipulation,” IEEE/ASME Transactions on Mechatronics , vol. 28, no. 3, pp. 1384–1394, 2023
2023
Earlier work this paper cites.
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” The International Journal of Robotics Research , p. 02783649241273668, 2023
2023
Earlier work this paper cites.
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar et al. , “Inner monologue: Embodied reasoning through planning with language models,” Conference on Robot Learning (CoRL) , 2023
2023
Earlier work this paper cites.
OpenAI, “Gpt-4 technical report,” 2023
2023
Earlier work this paper cites.
Y. Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y. J. Lee, and Y. Ma, “Investigating the catastrophic forgetting in multimodal large language models,” NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following. , 2023
2023
Earlier work this paper cites.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023
2023
Earlier work this paper cites.
L. Ren, J. Dong, S. Liu, L. Zhang, and L. Wang, “Embodied intelligence toward future smart manufacturing in the era of ai foundation model,” IEEE/ASME Transactions on Mechatronics , pp. 1–11, 2024
2024
Cited alongside, same era.
X. Li, M. Liu, H. Zhang, C. Yu, J. Xu, H. Wu, C. Cheang, Y. Jing, W. Zhang, H. Liu, H. Li, and T. Kong, “Vision-language foundation models as effective robot imitators,” The Twelfth International Conference on Learning Representations (ICLR) , 2024
2024
Cited alongside, same era.
P. Ding, H. Zhao, W. Zhang, W. Song, M. Zhang, S. Huang, N. Yang, and D. Wang, “Quar-vla: Vision-language-action model for quadruped robots,” in European Conference on Computer Vision . Springer, 2024, pp. 352–367
2024
Cited alongside, same era.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024
2024
Cited alongside, same era.
H. Hong, S. Wang, Z. Huang, Q. Wu, and J. Liu, “Navigating beyond instructions: Vision-and-language navigation in obstructed environments,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 7639–7648
2024
Later among the works it cites.
G. H. Chen, S. Chen, R. Zhang, J. Chen, X. Wu, Z. Zhang, Z. Chen, J. Li, X. Wan, and B. Wang, “Allava: Harnessing gpt4v-synthesized data for lite vision-language models,” 2024
2024
Later among the works it cites.
H. Zhao, M. Zhang, W. Zhao, P. Ding, S. Huang, and D. Wang, “Cobra: Extending mamba to multi-modal large language model for efficient inference,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 10, 2025, pp. 10 421–10 429
2025
Closest in time.
F. Tang, C. Liu, Z. Xu, M. Hu, Z. Huang, H. Xue, Z. Chen, Z. Peng, Z. Yang, S. Zhou et al. , “Seeing far and clearly: Mitigating hallucinations in mllms with attention causal decoding,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 26 147–26 159
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia, “Lisa: Reasoning segmentation via large language model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024, pp. 9579–9589
2024
Cited alongside, same era.
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang et al. , “Deepseek-vl: Towards real-world vision-language understanding,” CoRR , 2024
2024
Cited alongside, same era.
W. Song, H. Zhao, P. Ding, C. Cui, S. Lyu, Y. Fan, and D. Wang, “Germ: A generalist robotic model with mixture-of-experts for quadruped robot,” in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2024, pp. 11 879–11 886
2024
Cited alongside, same era.
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y. L. Tan, P. R. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine, “Octo: An open-source generalist robot policy,” Robotics: Science and Systems , 2024
2024
Cited alongside, same era.
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu, “3d diffusion policy,” Robotics: Science and Systems , 2024
2024
Cited alongside, same era.
A. Prasad, K. Lin, J. Wu, L. Zhou, and J. Bohg, “Consistency policy: Accelerated visuomotor policies via consistency distillation,” Robotics: Science and Systems , 2024
2024
Cited alongside, same era.
J. Gu, S. Kirmani, P. Wohlhart, Y. Lu, M. G. Arenas, K. Rao, W. Yu, C. Fu, K. Gopalakrishnan, Z. Xu, P. Sundaresan, P. Xu, H. Su, K. Hausman, C. Finn, Q. Vuong, and T. Xiao, “RT-trajectory: Robotic task generalization via hindsight trajectory sketches,” in The Twelfth International Conference on Learning Representations (ICLR) , 2024
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Q. Liang, W. Xiao, J. Long, and D. Zhang, “A multimodal robust recognition method for grasping objects with robot flexible grippers,” IEEE/ASME Transactions on Mechatronics , vol. 30, no. 2, pp. 1154–1165, 2025
2025
Closest in time.
Y. Han, K. Yu, R. Batra, N. Boyd, C. Mehta, T. Zhao, Y. She, S. Hutchinson, and Y. Zhao, “Learning generalizable vision-tactile robotic grasping strategy for deformable objects via transformer,” IEEE/ASME Transactions on Mechatronics , vol. 30, no. 1, pp. 554–566, 2025
2025
Closest in time.
J. Huang, K. Chen, J. Zhou, X. Lin, P. Abbeel, Q. Dou, and Y. Liu, “Dih-tele: Dexterous in-hand teleoperation framework for learning multiobjects manipulation with tactile sensing,” IEEE/ASME Transactions on Mechatronics , pp. 1–12, 2025
2025
Closest in time.
C. Zhou, R. Jiang, F. Luan, S. Meng, Z. Wang, Y. Dong, Y. Zhou, and B. He, “Dual-arm robotic fabric manipulation with quasi-static and dynamic primitives for rapid garment flattening,” IEEE/ASME Transactions on Mechatronics , pp. 1–11, 2025
2025
Closest in time.
Q. Bu, H. Li, L. Chen, J. Cai, J. Zeng, H. Cui, M. Yao, and Y. Qiao, “Towards synergistic, generalized, and efficient dual-system for robotic manipulation,” 2025
2025
Closest in time.
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, Y. Fang, D. Fox, F. Hu, S. Huang, J. Jang et al. , “Gr00t n1: An open foundation model for generalist humanoid robots,” Arxiv , 2025
2025
Closest in time.
2025
Closest in time.
Y. Zhou, Y. Zhou, K. Jin, and H. Wang, “Hierarchical reinforcement learning with model guidance for mobile manipulation,” IEEE/ASME Transactions on Mechatronics , pp. 1–9, 2025
2025
Closest in time.