Fetching the paper…
Reading the bibliography…
Vision-language-action models have emerged as a crucial paradigm in robotic manipulation.
The graph neural network model
Scarselli, F.; Gori, M.; Tsoi, A. C.; Hagenbuchner, M.; and Monfardini, G. 2008 · 2008
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Dabis, J.; Finn, C.; Gopalakrishnan, K.; Hausman, K.; Herzog, A.; Hsu, J.; et al. 2022 · 2022
Earlier work this paper cites.
Flow matching for generative modeling
Lipman, Y.; Chen, R. T.; Ben-Hamu, H.; Nickel, M.; and Le, M. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; and Song, S. 2023 · 2023
Earlier work this paper cites.
Cross-entropy loss functions: Theoretical analysis and applications
Mao, A.; Mohri, M.; and Zhong, Y. 2023 · 2023
Earlier work this paper cites.
Unleashing large-scale video generative pre-training for visual robot manipulation
Wu, H.; Jing, Y.; Cheang, C.; Chen, G.; Xu, J.; Li, X.; Liu, M.; Li, H.; and Kong, T. 2023 · 2023
Earlier work this paper cites.
Crossformer: Transformer utilizing cross-dimension dependency for multivariate time series forecasting
Zhang, Y.; and Yan, J. 2023 · 2023
Earlier work this paper cites.
Learning fine-grained bimanual manipulation with low-cost hardware
Zhao, T. Z.; Kumar, V.; Levine, S.; and Finn, C. 2023 · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Zitkovich, B.; Yu, T.; Xu, S.; Xu, P.; Xiao, T.; Xia, F.; Wu, J.; Wohlhart, P.; Welker, S.; Wahid, A.; et al. 2023 · 2023
Cited alongside, same era.
Paligemma: A versatile 3b vlm for transfer
Beyer, L.; Steiner, A.; Pinto, A. S.; Kolesnikov, A.; Wang, X.; Salz, D.; Neumann, M.; Alabdulmohsin, I.; Tschannen, M.; Bugliarello, E.; et al. 2024 · 2024
Cited alongside, same era.
p i _ 0 pi\_0 : A Vision-Language-Action Flow Model for General Robot Control
Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; et al. 2024 · 2024
Cited alongside, same era.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
Octo: An open-source generalist robot policy
Team, O. M.; Ghosh, D.; Walke, H.; Pertsch, K.; Black, K.; Mees, O.; Dasari, S.; Hejna, J.; Kreiman, T.; Xu, C.; et al. 2024 · 2024
Later among the works it cites.
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
Wen, J.; Zhu, M.; Zhu, Y.; Tang, Z.; Li, J.; Zhou, Z.; Li, C.; Liu, X.; Peng, Y.; Shen, C.; et al. 2024 · 2024
Later among the works it cites.
Robotic control via embodied chain-of-thought reasoning
Zawalski, M.; Chen, W.; Pertsch, K.; Mees, O.; Finn, C.; and Levine, S. 2024 · 2024
Later among the works it cites.
Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; et al. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cheang, C.-L.; Chen, G.; Jing, Y.; Kong, T.; Li, H.; Li, Y.; Liu, Y.; Wu, H.; Xu, J.; Yang, Y.; et al. 2024 · 2024
Cited alongside, same era.
Yolo-world: Real-time open-vocabulary object detection
Cheng, T.; Song, L.; Ge, Y.; Liu, W.; Wang, X.; and Shan, Y. 2024 · 2024
Cited alongside, same era.
Openvla: An open-source vision-language-action model
Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.; Lam, G.; Sanketi, P.; et al. 2024 · 2024
Cited alongside, same era.
Rdt-1b: a diffusion foundation model for bimanual manipulation
Liu, S.; Wu, L.; Li, B.; Tan, H.; Chen, H.; Wang, Z.; Xu, K.; Su, H.; and Zhu, J. 2024 · 2024
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
O’Neill, A.; Rehman, A.; Maddukuri, A.; Gupta, A.; Padalkar, A.; Lee, A.; Pooley, A.; Gupta, A.; Mandlekar, A.; Jain, A.; et al. 2024 · 2024
Cited alongside, same era.
Dexvla: Vision-language model with plug-in diffusion expert for general robot control
Wen, J.; Zhu, Y.; Li, J.; Tang, Z.; Shen, C.; and Feng, F. 2025a
Cited in the paper.
Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation
Wen, J.; Zhu, Y.; Li, J.; Zhu, M.; Tang, Z.; Wu, K.; Xu, Z.; Liu, N.; Cheng, R.; Shen, C.; et al. 2025b
Cited in the paper.
Chen, W.; Belkhale, S.; Mirchandani, S.; Mees, O.; Driess, D.; Pertsch, K.; and Levine, S. 2025 · 2025
Closest in time.
Vision Language Action Models in Robotic Manipulation: A Systematic Review
Din, M. U.; Akram, W.; Saoud, L. S.; Rosell, J.; and Hussain, I. 2025 · 2025
Closest in time.
Spatialvla: Exploring spatial representations for visual-language-action model
Qu, D.; Song, H.; Chen, Q.; Yao, Y.; Ye, X.; Ding, Y.; Wang, Z.; Gu, J.; Zhao, B.; Wang, D.; et al. 2025 · 2025
Closest in time.
Wang, S. 2025 · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models
Zhao, Q.; Lu, Y.; Kim, M. J.; Fu, Z.; Zhang, Z.; Wu, Y.; Li, Z.; Ma, Q.; Han, S.; Finn, C.; et al. 2025 · 2025
Closest in time.