Fetching the paper…
Reading the bibliography…
Robotic manipulation is a fundamental component of automation.
2022
Earlier work this paper cites.
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid et al. , “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” in Conference on Robot Learning . PMLR, 2023, pp. 2165–2183
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in Proceedings of the IEEE/CVF international conference on computer vision , 2023, pp. 11 975–11 986
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in Proceedings of the IEEE/CVF international conference on computer vision , 2023, pp. 4195–4205
2023
Earlier work this paper cites.
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du et al. , “Bridgedata v2: A dataset for robot learning at scale,” in Conference on Robot Learning . PMLR, 2023, pp. 1723–1736
2023
Earlier work this paper cites.
B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone, “Libero: Benchmarking knowledge transfer for lifelong robot learning,” Advances in Neural Information Processing Systems , vol. 36, pp. 44 776–44 791, 2023
2023
Earlier work this paper cites.
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” The International Journal of Robotics Research , p. 02783649241273668, 2023
2023
Earlier work this paper cites.
Y. Zhang, T. Xue, A. Razmjoo, and S. Calinon, “Logic learning from demonstrations for multi-step manipulation tasks in dynamic environments,” IEEE Robotics and Automation Letters , vol. 9, no. 8, pp. 7214–7221, 2024
2024
Earlier work this paper cites.
Y. Yang, Z. Cui, Q. Zhang, and J. Liu, “Ps6d: Point cloud based symmetry-aware 6d object pose estimation in robot bin-picking,” in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2024, pp. 7167–7174
2024
Earlier work this paper cites.
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain et al. , “Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 6892–6903
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
G. Wang, H. Li, S. Zhang, D. Guo, Y. Liu, and H. Liu, “Observe then act: Asynchronous active vision-action model for robotic manipulation,” IEEE Robotics and Automation Letters , vol. 10, no. 4, pp. 3422–3429, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
L. Xiaoyang Shi, Z. Hu, T. Z. Zhao, A. Sharma, K. Pertsch, J. Luo, S. Levine, and C. Finn, “Yell at your robot: Improving on-the-fly from language corrections,” arXiv e-prints , pp. arXiv–2403, 2024
2024
Cited alongside, same era.
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh, “Prismatic vlms: Investigating the design space of visually-conditioned language models,” in Forty-first International Conference on Machine Learning , 2024
2024
Cited alongside, same era.
X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao, “Evaluating real-world robot manipulation policies in simulation,” in Proceedings of the Conference on Robot Learning (CoRL) , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
W. Xia, R. Feng, D. Wang, and D. Hu, “Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 6981–6990
2025
Closest in time.
2025
Closest in time.
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn et al. , “Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 1702–1713
2025
Closest in time.
2025
Closest in time.