Fetching the paper…
Reading the bibliography…
Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general-purpose embodied intelligence.
A. Kembhavi, M. Salvato, E. Kolve, et al. , “A diagram is worth a dozen images,” in European conference on computer vision . Springer, 2016, pp. 235–251
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, et al. , “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, et al. , “Flamingo: a visual language model for few-shot learning,” Advances in neural information processing systems , vol. 35, pp. 23 716–23 736, 2022
2022
Earlier work this paper cites.
H. Liu, C. Li, Q. Wu, et al. , “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, pp. 34 892–34 916, 2023
2023
Earlier work this paper cites.
A. Brohan, N. Brown, J. Carbajal, et al. , “Rt-1: Robotics Transformer for Real-World Control at Scale.” in Robotics: Science and Systems Conference (RSS) , 2023
2023
Earlier work this paper cites.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in Proceedings of the IEEE/CVF international conference on computer vision , 2023, pp. 4195–4205
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
M. Zawalski, W. Chen, K. Pertsch, et al. , “Robotic control via embodied chain-of-thought reasoning,” in 8th Annual Conference on Robot Learning , 2024
2024
Earlier work this paper cites.
OpenAI, :, A. Hurst, et al. , “Gpt-4o system card,” 2024
2024
Earlier work this paper cites.
B. Xiao, H. Wu, W. Xu, et al. , “Florence-2: Advancing a unified representation for a variety of vision tasks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 4818–4829
2024
Cited alongside, same era.
X. Yue, Y. Ni, K. Zhang, et al. , “Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 9556–9567
2024
Cited alongside, same era.
L. Chen, J. Li, X. Dong, et al. , “Are we on the right way for evaluating large vision-language models?” Advances in Neural Information Processing Systems , vol. 37, pp. 27 056–27 087, 2024
2024
Cited alongside, same era.
C. Fu, P. Chen, Y. Shen, et al. , “Mme: A comprehensive evaluation benchmark for multimodal large language models,” 2024
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
J. Wen, Y. Zhu, J. Li, et al. , “Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation,” IEEE Robotics and Automation Letters , 2025
2025
Closest in time.
J. Liu, H. Chen, P. An, et al. , “Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model,” 2025
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Guan, F. Liu, X. Wu, et al. , “Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 14 375–14 385
2024
Cited alongside, same era.
R. Team, “Realworldqa,” https://x.ai/news/grok-1.5v, 2024
2024
Cited alongside, same era.
H. Duan, J. Yang, Y. Qiao, et al. , “Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,” in Proceedings of the 32nd ACM international conference on multimedia , 2024, pp. 11 198–11 201
2024
Cited alongside, same era.
2025
Cited alongside, same era.
M. J. Kim, K. Pertsch, S. Karamcheti, et al. , “Openvla: An open-source vision-language-action model,” in Proceedings of The 8th Conference on Robot Learning , ser. Proceedings of Machine Learning Research, P. Agrawal, O. Kroemer, and W. Burgard, Eds., vol. 270. PMLR, 06–09 Nov 2025, pp. 2679–2713
2025
Cited alongside, same era.
2025
Cited alongside, same era.
2025
Closest in time.
Z. Zhou, Y. Zhu, M. Zhu, et al. , “Chatvla: Unified multimodal understanding and robot control with vision-language-action model,” 2025
2025
Closest in time.
P. Intelligence, K. Black, N. Brown, et al. , “ π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization,” 2025
2025
Closest in time.
C. Cheang, S. Chen, Z. Cui, et al. , “Gr-3 technical report,” 2025
2025
Closest in time.
J. Yang, R. Tan, Q. Wu, et al. , “Magma: A foundation model for multimodal ai agents,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2025, pp. 14 203–14 214
2025
Closest in time.
Q. Sun, P. Hong, T. D. Pala, et al. , “Emma-X: An embodied multimodal action model with grounded chain of thought and look-ahead spatial reasoning,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , July 2025, pp. 14 199–14 214
2025
Closest in time.