Fetching the paper…
Reading the bibliography…
Robotic manipulation in 3D requires effective computation of N degree-of-freedom joint-space trajectories that enable precise and robust control.
C. R. Qi, H. Su, et al., “PointNet: Deep learning on point sets for 3D classification and segmentation,” IEEE/CVF Computer Vision and Pattern Recognition , 2017
2017
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy et al., “Learning transferable visual models from natural language supervision,” 38th International Conference on Machine Learning , 2021
2021
Earlier work this paper cites.
M. Shridhar, L. Manuelli and D. Fox, “CLIPort: What and where pathways for robotic manipulation,” The 5th Conference on Robot Learning , 2022
2022
Earlier work this paper cites.
A. Brohan, N. Brown, J. Carbajal, et al., “RT-1: Robotics transformer for real-world control at scale,” Robotics: Science and Systems , 2023
2023
Earlier work this paper cites.
A. Brohan, N. Brown, J. Carbajal, et al., “RT-2: Vision-language-action models transfer web knowledge to robotic control,” The 7th Conference on Robot Learning , 2023
2023
Earlier work this paper cites.
C. Chi, Z. Xu, et al., “Diffusion policy: Visuomotor policy learning via action diffusion,” International Journal of Robotics Research , 2023
2023
Earlier work this paper cites.
C. Huang, O. Mees, A. Zeng, et al., “Visual language maps for robot navigation,” IEEE International Conference on Robotics and Automation , 2023
2023
Earlier work this paper cites.
A. Z. Ren, A. Dixit, A. Bodrova et al., “Robots that ask for help: Uncertainty alignment for large language model planners,” The 6th Conference on Robot Learning , 2023
2023
Earlier work this paper cites.
I. Singh, V. Blukis, A. Mousavian et al., “ProgPrompt: Generating situated robot task plans using large language models,” IEEE International Conference on Robotics and Automation , 2023
2023
Earlier work this paper cites.
C. Song, J. Wu, C. Washington et al., “LLM-Planner: Few-shot grounded planning for embodied agents with large language models,” IEEE/CVF International Conference on Computer Vision , 2023
2023
Earlier work this paper cites.
Z. J. Cui, H. Pan, A. Iyer, et al., “Dynamo: In-domain dynamics pretraining for visuo-motor control,” Advances in Neural Information Processing Systems , 2024
2024
Earlier work this paper cites.
M. Deitke, C. Clark, S. Lee, et al., “Molmo and Pixmo: Open weights and open data for state-of-the-art multimodal models,” IEEE/CVF Computer Vision and Pattern Recognition , 2024
2024
Earlier work this paper cites.
Q. Gu, A. Kuwajerwala, S. Morin, et al., “ConceptGraphs: Open-vocabulary 3D scene graphs for perception and planning,” IEEE International Conference on Robotics and Automation , 2024
2024
Earlier work this paper cites.
A. O’Neill, A. Rehman, Abhinav Gupta et al., “Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration,” IEEE International Conference on Robotics and Automation , 2024
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
U. Zaratiana, N. Tomeh, P. Holat, and T. Charnois, “GLiNER: Generalist model for named entity recognition using bidirectional transformer,” Conference of the North American Chapter of the Association for Computational Linguistics , 2024
2024
Cited alongside, same era.
M. Zawalski, W. Chen, K. Pertsch et al., “Robotic control via embodied chain-of-thought reasoning,” The 8th Conference on Robot Learning , 2024
2024
Cited alongside, same era.
J. Yang, H. Zhu, et al., “Tra-MoE: Learning trajectory prediction model from multiple domains for adaptive policy conditioning,” IEEE/CVF Computer Vision and Pattern Recognition , 2025
2025
Closest in time.
G. Zhi, Z. Zhang et al., “Closed-loop open-vocabulary mobile manipulation with GPT-4V,” IEEE International Conference on Robotics and Automation , 2025
2025
Closest in time.
K. Black, N. Brown, D. Driess et al., ” π 0 \pi_{0} : A Vision-Language-Action Flow Model for General Robot Control”, Robotics: Science and Systems , 2025
2025
Closest in time.
Q. Zhao, Y. Lu, M. Kim, et al., “CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models,” IEEE/CVF Computer Vision and Pattern Recognition , 2025
2025
Closest in time.
D. Qu, H. Song, Q. Chen, et al., “SpatialVLA Exploring Spatial Representations for Visual-Language-Action Models,” Robotics: Science and Systems , 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Ghosh, H. Walke, K. Pertsch et al., “Octo: An Open-Source Generalist Robot Policy,” Robotics: Science and Systems , 2024
2024
Cited alongside, same era.
M. Reuss, O. Yagmurlu, F. Wenzel et al. ”Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals”, Robotics: Science and Systems , 2024
2024
Cited alongside, same era.
C. Song, V. Blukis, et al. ”RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics,” IEEE/CVF Computer Vision and Pattern Recognition , 2025
2025
Cited alongside, same era.
H. Huang, F. Liu, L. Fu, et al., “Otter: A vision-language-action model with text-aware visual feature extraction,” 42nd International Conference on Machine Learning , 2025
2025
Cited alongside, same era.
G.-C. Kang, J. Kim, K. Shim, et al., “CLIP-RT: Learning language-conditioned robotic policies from natural language supervision,” Robotics: Science and Systems , 2025
2025
Cited alongside, same era.
M. J. Kim, K. Pertsch, S. Karamcheti, et al., “OpenVLA: An open-source vision-language-action model,” The 8th Conference on Robot Learning , 2025
2025
Cited alongside, same era.
M. J. Kim, C. Finn, and P. Liang, “Fine-tuning vision-language-action models: Optimizing speed and success,” Robotics: Science and Systems , 2025
2025
Cited alongside, same era.
P. Sundaresan, H. Hu, Q. Vuong, et al., “What’s the move? Hybrid imitation learning via salient points,” International Conference on Learning Representations , 2025
2025
Cited alongside, same era.
2025
Closest in time.
L. Jinming, Z. Yichen, T. Zhibin, et al., “CoA-VLA: Improving vision-language-action models via visual-textual chain-of-affordance,” IEEE/CVF International Conference on Computer Vision , 2025
2025
Closest in time.
W. Chen, S. Belkhale, S. Mirchandani et al., “Training Strategies for Efficient Embodied Reasoning,” The 9th Conference on Robot Learning , 2025
2025
Closest in time.
N. Ravi, V. Gabeur, Y. Hu, et al., “SAM 2: Segment Anything in Images and Videos,” International Conference on Learning Representations , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
R. Goswami, et al., “RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Traininge,” IEEE/CVF Computer Vision and Pattern Recognition , 2025
2025
Closest in time.
V. Bhat, S. Kim, V. Blukis, et al., “BOP-ASK: Object-Interaction Reasoning for Vision-Language Models,” IEEE/CVF Computer Vision and Pattern Recognition , 2026
2026
Closest in time.
V. Bhat, N. Patel, et al., ”MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping,” IEEE/CVF Winter Conference on Applications of Computer Vision , 2026
2026
Closest in time.