Fetching the paper…
Reading the bibliography…
The pursuit of out-of-distribution generalization in Vision-Language-Action (VLA) models is often hindered by catastrophic forgetting of the Vision-Language Model (VLM) backbone during fine-tuning.
S. Yenamandra, A. Ramachandran, K. Yadav, A. S. Wang, M. Khanna, T. Gervet, T.-Y. Yang, V. Jain, A. Clegg, J. M. Turner et al. , “Homerobot: Open-vocabulary mobile manipulation,” in Conference on Robot Learning . PMLR, 2023, pp. 1975–2011
2011
Earlier work this paper cites.
A. Singh, “Cmu 10-704: Information processing and learning,” https://www.cs.cmu.edu/~aarti/Class/10704/ , 2012
2012
Earlier work this paper cites.
M. C. Koval, N. S. Pollard, and S. S. Srinivasa, “Pre-and post-contact policy decomposition for planar contact manipulation under uncertainty,” The International Journal of Robotics Research , vol. 35, no. 1-3, pp. 244–264, 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
Y. Zhou, C. Barnes, J. Lu, J. Yang, and H. Li, “On the continuity of rotation representations in neural networks,” in CVPR , 2019, pp. 5738–5746
2019
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
K. Xu, H. Yu, Q. Lai, Y. Wang, and R. Xiong, “Efficient learning of goal-oriented push-grasping synergy in clutter,” IEEE Robotics and Automation Letters , vol. 6, no. 4, pp. 6337–6344, 2021
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog et al. , “Do as i can, not as i say: Grounding language in robotic affordances,” in Conference on Robot Learning , 2022
2022
Earlier work this paper cites.
K. Xu, H. Yu, R. Huang, D. Guo, Y. Wang, and R. Xiong, “Efficient object manipulation to an arbitrary goal pose: Learning-based anytime prioritized planning,” in 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 2022, pp. 7277–7283
2022
Earlier work this paper cites.
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn, “Bc-z: Zero-shot task generalization with robotic imitation learning,” in Conference on Robot Learning . PMLR, 2022, pp. 991–1002
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al. , “Lora: Low-rank adaptation of large language models.” ICLR , vol. 1, no. 2, p. 3, 2022
2022
Earlier work this paper cites.
C. Zhou, C. C. Loy, and B. Dai, “Extract free dense labels from clip,” in European Conference on Computer Vision , 2022, pp. 696–712
2022
Earlier work this paper cites.
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid et al. , “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” in Conference on Robot Learning . PMLR, 2023, pp. 2165–2183
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone, “Libero: Benchmarking knowledge transfer for lifelong robot learning,” Advances in Neural Information Processing Systems , vol. 36, pp. 44 776–44 791, 2023
2023
Earlier work this paper cites.
H.-S. Fang, C. Wang, H. Fang, M. Gou, J. Liu, H. Yan, W. Liu, Y. Xie, and C. Lu, “Anygrasp: Robust and efficient grasp perception in spatial and temporal domains,” IEEE Transactions on Robotics , vol. 39, no. 5, pp. 3929–3945, 2023
2023
Earlier work this paper cites.
R. Firoozi, J. Tucker, S. Tian, A. Majumdar, J. Sun, W. Liu, Y. Zhu, S. Song, A. Kapoor, K. Hausman et al. , “Foundation models in robotics: Applications, challenges, and the future,” The International Journal of Robotics Research , p. 02783649241281508, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
K. Xu, S. Zhao, Z. Zhou, Z. Li, H. Pi, Y. Zhu, Y. Wang, and R. Xiong, “A joint modeling of vision-language-action for target-oriented grasping in clutter,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 11 597–11 604
2023
Earlier work this paper cites.
Y. Zhu, A. Joshi, P. Stone, and Y. Zhu, “Viola: Imitation learning for vision-based manipulation with object proposal priors,” in Conference on Robot Learning . PMLR, 2023, pp. 1199–1210
2023
Earlier work this paper cites.
A. Rashid, S. Sharma, C. M. Kim, J. Kerr, L. Y. Chen, A. Kanazawa, and K. Goldberg, “Language embedded radiance fields for zero-shot task-oriented grasping,” in 7th Annual Conference on Robot Learning , 2023
2023
Earlier work this paper cites.
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng, “Code as policies: Language model programs for embodied control,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 9493–9500
2023
Earlier work this paper cites.
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei, “Voxposer: Composable 3d value maps for robotic manipulation with language models,” in Conference on Robot Learning . PMLR, 2023, pp. 540–562
2023
Earlier work this paper cites.
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, Y. Dou, Y. Chen, L. Fei-Fei, A. Anandkumar, Y. Zhu, and L. Fan, “Vima: robot manipulation with multimodal prompts,” in Proceedings of the 40th International Conference on Machine Learning , 2023, pp. 14 975–15 022
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” in Proceedings of Robotics: Science and Systems (RSS) , 2023
2023
Earlier work this paper cites.
J. Anschütz and A.-C. Le Bras, “Prismatic dieudonné theory,” in Forum of Mathematics, Pi , vol. 11. Cambridge University Press, 2023, p. e2
2023
Earlier work this paper cites.
M. Mittal, C. Yu, Q. Yu, J. Liu, N. Rudin, D. Hoeller, J. L. Yuan, R. Singh, Y. Guo, H. Mazhar, A. Mandlekar, B. Babich, G. State, M. Hutter, and A. Garg, “Orbit - A Unified Simulation Framework for Interactive Robot Learning Environments,” IEEE Robotics and Automation Letters , vol. 8, no. 6, 2023
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn, “OpenVLA: An open-source vision-language-action model,” in 8th Annual Conference on Robot Learning , 2024. [Online]. Available: https://openreview.net/forum?id=ZMnD6QZAE6
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
K. Xu, X. Xia, K. Wang, Y. Yang, Y. Mao, B. Deng, J. Ye, R. Xiong, and Y. Wang, “Efficient alignment of unconditioned action prior for language-conditioned pick and place in clutter,” IEEE Transactions on Automation Science and Engineering , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
M. Zhu, Y. Zhu, J. Li, J. Wen, Z. Xu, N. Liu, R. Cheng, C. Shen, Y. Peng, F. Feng et al. , “Scaling diffusion policy in transformer to 1 billion parameters for robotic manipulation,” in 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2025, pp. 10 838–10 845
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2024
Cited alongside, same era.
K. Xu, Z. Zhou, J. Wu, H. Lu, R. Xiong, and Y. Wang, “Grasp, see and place: Efficient unknown object rearrangement with policy structure prior,” IEEE Transactions on Robotics , 2024
2024
Cited alongside, same era.
Y. Wang, M. Zhang, Z. Li, T. Kelestemur, K. R. Driggs-Campbell, J. Wu, L. Fei-Fei, and Y. Li, “D 3 fields: Dynamic 3d descriptor fields for zero-shot generalizable rearrangement,” in 8th Annual Conference on Robot Learning , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki, “3d diffuser actor: Policy diffusion with 3d scene representations,” in 8th Annual Conference on Robot Learning , 2024
2024
Cited alongside, same era.
A. Goyal, V. Blukis, J. Xu, Y. Guo, Y.-W. Chao, and D. Fox, “Rvt-2: Learning precise manipulation from few demonstrations,” in Robotics: Science and Systems (RSS) , 2024
2024
Cited alongside, same era.
A. D. Vuong, M. N. Vu, H. Le, B. Huang, H. T. T. Binh, T. Vo, A. Kugi, and A. Nguyen, “Grasp-anything: Large-scale grasp dataset from foundation models,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 14 030–14 037
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
G. A. Team, “Gen-0: Embodied foundation models that scale with physical interaction,” Generalist AI Blog , 2025, https://generalistai.com/blog/preview-uqlxvb-bb.html
2025
Closest in time.
H. Liu, X. Li, P. Li, M. Liu, D. Wang, J. Liu, B. Kang, X. Ma, T. Kong, and H. Zhang, “Towards generalist robot policies: What matters in building vision-language-action models,” 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn et al. , “Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 1702–1713
2025
Closest in time.
2025
Closest in time.
R. Team, “Rdt2: Enabling zero-shot cross-embodiment generalization by scaling up umi data,” September 2025. [Online]. Available: https://github.com/thu-ml/RDT2
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang et al. , “Gr00t n1.5: An improved open foundation model for generalist humanoid robots,” 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.