Fetching the paper…
Reading the bibliography…
Vision Language Action (VLA) models represent a transformative shift in robotics, with the aim of unifying visual perception, natural language understanding, and embodied control within a single learning framework.
Language models are few-shot learners
Brown, T.B., Mann, B., Ryder, N., et al., 2020 · 1901
Earlier work this paper cites.
Thomason, J., Murray, M., Cakmak, M., Zettlemoyer, L., 2019 · 1907
Earlier work this paper cites.
Shridhar, M., Thomason, J., Gordon, D., Bisk, Y., Han, W., Mottaghi, R., Zettlemoyer, L., Fox, D., 2020 · 1912
Earlier work this paper cites.
Covla: Comprehensive vision-language-action dataset for autonomous driving
Arai, H., Miwa, K., Sasaki, K., Watanabe, K., Yamaguchi, Y., Aoki, S., Yamamoto, I., 2025 · 1943
Earlier work this paper cites.
Design and use paradigms for gazebo, an open-source multi-robot simulator, in: 2004 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE. pp. 2149–2154
Koenig, N., Howard, A., 2004 · 2004
Earlier work this paper cites.
Krantz, J., Wijmans, E., Mukhopadhyay, A., Lee, S., Chernova, S., Batra, D., 2020 · 2004
Earlier work this paper cites.
Webots: Professional mobile robot simulation, in: 2004 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE. pp. 4020–4025
Michel, O., 2004 · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L., 2009 · 2009
Earlier work this paper cites.
robosuite: A modular simulation framework and benchmark for robot learning
Zhu, Y., Gupta, A., Ebert, F., et al., 2020 · 2009
Earlier work this paper cites.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N., 2021 · 2010
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., et al., 2020 · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control, in: 2012 IEEE/RSJ international conference on intelligent robots and systems, IEEE. pp. 5026–5033
Todorov, E., Erez, T., Tassa, Y., 2012 · 2012
Earlier work this paper cites.
V-rep: A versatile and scalable robot simulation framework, in: 2013 IEEE/RSJ international conference on intelligent robots and systems, IEEE. pp. 1321–1326
Rohmer, E., Singh, S.P., Freese, M., 2013 · 2013
Earlier work this paper cites.
Pybullet, a python module for physics simulation for robotics, games and machine learning
Coumans, E., Bai, Y., 2016 · 2016
Earlier work this paper cites.
AI2-THOR: An interactive 3d environment for visual ai, in: Proceedings of the 1st Annual Conference on Robot Learning (CoRL)
Kolve, E., Mottaghi, R., Han, W., Randhavane, T., Zheng, X., Li, Y., Gupta, A., Farhadi, A., 2017 · 2017
Earlier work this paper cites.
Press, O., Wolf, L., 2017 · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Å., Polosukhin, I., 2017 · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., van den Hengel, A., 2018 · 2018
Earlier work this paper cites.
Embodied question answering, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., Batra, D., 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2018 · 2018
Earlier work this paper cites.
Unity: A general platform for intelligent agents, in: Proceedings of the 1st Annual Conference on Robot Learning (CoRL), pp. 49–60
Juliani, A., Berges, V., Teng, E., Gao, Y., Henry, H., Mattar, M., Lange, D., 2018 · 2018
Earlier work this paper cites.
Habitat: A platform for embodied ai research
Savva, M., Chang, A.X., Dosovitskiy, A., et al., 2019 · 2019
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
James, S., Ma, Z., Arrojo, D.R., Davison, A.J., 2020 · 2020
Earlier work this paper cites.
Adapterfusion: Non-destructive task composition for transfer learning, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 4875–4884
Pfeiffer, J., Vulic, I., Gurevych, I., 2020 · 2020
Earlier work this paper cites.
Xia, F., Li, C., Martín-Martín, R., Litany, O., Zamir, A.R., Savarese, S., 2020 · 2020
Earlier work this paper cites.
Sapien: A simulated part-based interactive environment, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11097–11107
Xiang, F., Qin, Y., Mo, K., Xia, Y., Zhu, H., Liu, F., Liu, M., Jiang, H., Yuan, Y., Wang, H., et al., 2020 · 2020
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning, in: Conference on robot learning, PMLR. pp. 1094–1100
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., Levine, S., 2020 · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers, in: ICCV
Caron, M., Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., Jégou, H., 2021 · 2021
Earlier work this paper cites.
Pre-train or prompt? exploring the encoder-decoder framework for zero-shot learning
Liu, X., Chen, W., Chen, Y., Chen, Y.S., Wang, W.Y., 2021 · 2021
Earlier work this paper cites.
Isaac gym: High performance gpu based physics simulation for robot learning
Makoviychuk, V., Wawrzyniak, L., Rathod, Y., Allshire, A., Handa, A., Müller, J., Widmaier, F., Leal-Taixé, L., Makadia, A., Leutenegger, S., 2021 · 2021
Earlier work this paper cites.
Teach: Task-driven embodied agents that chat
Padmakumar, A., Thomason, J., Shrivastava, A., Lange, P., Narayan-Chen, A., Gella, S., Piramuthu, R., Tur, G., Hakkani-Tur, D., 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J.W., Hallacy, C., et al., 2021 · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., et al., 2022 · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al., 2022 · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H.W., Sutton, C., Sepassi, R., Gehrmann, S., Elsen, E., Patrick, D., Mishkin, P., 2022 · 2022
Earlier work this paper cites.
DialFRED: Dialogue-Enabled Agents for Embodied Instruction Following
Gao, X., Gao, Q., Gong, R., Lin, K., Thattai, G., Sukhatme, G.S., 2022a · 2022
Earlier work this paper cites.
Dialfred: Dialogue-enabled agents for embodied instruction following
Gao, X., Gao, Q., Gong, R., Lin, K., Thattai, G., Sukhatme, G.S., 2022b · 2022
Earlier work this paper cites.
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., et al., R.G., 2022 · 2022
Earlier work this paper cites.
Perceiver IO: A general architecture for structured inputs & outputs, in: International Conference on Learning Representations (ICLR)
Jaegle, A., Gimeno, N., Brock, A., Zisserman, A., Carreira, J., Vinyals, O., Verdegaal, R., Pessoa, P., Nowozin, S., 2022 · 2022
Earlier work this paper cites.
Vima: General robot manipulation with multimodal prompts
Jiang, Y., Gupta, A., Zhang, Z., et al., 2022 · 2022
Earlier work this paper cites.
Align before fuse: Vision and language representation learning with momentum distillation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18713–18723
Li, J., Li, X., Li, X., Huang, J., Zhang, J., Wang, L., Dou, Q., Ling, H., 2022 · 2022
Earlier work this paper cites.
Mees, O., Hermann, L., Rosete-Beas, E., Burgard, W., 2022 · 2022
Earlier work this paper cites.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J., et al., 2022 · 2022
Earlier work this paper cites.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts, in: ACM International Conference on Multimedia (MM)
Wang, L., Zhang, H., Zhao, Y., Liu, Z., Bian, J., Yu, H., Xu, C., Lau, R., Wang, S., 2022 · 2022
Earlier work this paper cites.
Roboagent: Generalist robot agent with semantic and temporal understanding
Bharadhwaj, H., Pore, N., Liang, J., Singh, J., Rao, K., Zeng, A., Gopalakrishnan, K., 2023 · 2023
Earlier work this paper cites.
Seq2code: Encoder-decoder model for program synthesis
Brandišauskas, M., Žukauskas, M., Krizhanovsky, A., 2023 · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D., Ruiz, N., Goyal, K., Chebotar, Y., Irpan, A., Ailon, X., Levine, S., Finn, C., 2023 · 2023
Cited alongside, same era.
Robotic task generalization via hindsight trajectory sketches
Gu, J., Kirmani, S., Wohlhart, P., Lu, Y., Arenas, M.G., Rao, K., Yu, W., Fu, C., Gopalakrishnan, K., Xu, Z., et al., 2023 · 2023
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models
Huang, W., Wang, C., Zhang, R., Li, Y., Wu, J., Fei-Fei, L., 2023 · 2023
Cited alongside, same era.
Deep learning with vision transformers: A survey
Lam, C., Wang, X., Lu, X., Yao, Y., Yang, M.H., 2023 · 2023
π \pi -0.5:: A vision-language-action model with open-world generalization
Black, K., Brown, N., Darpinian, J., Dhabalia, K., Driess, D., Esmail, A., Equi, M., et al., 2025 · 2025
Closest in time.
Bu, Q., Cai, J., Chen, L., Cui, X., Ding, Y., Feng, S., Gao, S., He, X., Hu, X., Huang, X., et al., 2025 · 2025
Closest in time.
Open x-embodiment: Robotic learning datasets and rt-x models
Collaboration, E., O’Neill, A., Rehman, A., Gupta, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., et al., A.P., 2025 · 2025
Closest in time.
Humanoid-vla: Towards universal humanoid control with visual integration
Ding, P., Ma, J., Tong, X., Zou, B., Luo, X., Fan, Y., Wang, T., Lu, H., Mo, P., Liu, J., et al., 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 10965–10975
Li, J., Peng, E.A., Wang, C., Liu, J., Feichtenhofer, C., 2023 · 2023
Cited alongside, same era.
Robo360: A 3d omnispective multi-material robotic manipulation dataset
Liang, L., Bian, L., Xiao, C., et al., 2023 · 2023
Cited alongside, same era.
Libero: Benchmarking knowledge transfer for lifelong robot learning
Liu, B., Zhu, Y., Gao, C., Feng, Y., Liu, Q., Zhu, Y., Stone, P., 2023 · 2023
Cited alongside, same era.
Perceiver-actor: A multi-task transformer for robotic manipulation, in: Conference on Robot Learning, PMLR. pp. 785–799
Shridhar, M., Manuelli, L., Fox, D., 2023 · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Martin, T., Stone, L., Albert, A., Almahairi, A., Laradji, I., Aqaj, Y., Baratin, A., Lee, S., Verde, Z., Kaplanyan, A., Azar, M., Gelly, S., Joulin, A., 2023 · 2023
Cited alongside, same era.
Bridgedata v2: A dataset for robot learning at scale, in: Conference on Robot Learning, PMLR. pp. 1723–1736
Walke, H.R., Black, K., Zhao, T.Z., Vuong, Q., Zheng, C., Hansen-Estruch, P., He, A.W., Myers, V., Kim, M.J., Du, M., et al., 2023 · 2023
Cited alongside, same era.
Dexgraspnet: A large-scale robotic dexterous grasp dataset for general objects based on simulation, in: 2023 IEEE International Conference on Robotics and Automation (ICRA), IEEE. pp. 11359–11366
Wang, R., Zhang, J., Chen, J., Xu, Y., Li, P., Liu, T., Wang, H., 2023 · 2023
Cited alongside, same era.
Driess, D., Springenberg, J.T., Ichter, B., Yu, L., Li-Bell, A., Pertsch, K., Ren, A.Z., Walke, H., Vuong, Q., Shi, L.X., et al., 2025 · 2025
Closest in time.
Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions
Fan, C., Jia, X., Sun, Y., Wang, Y., Wei, J., Gong, Z., Zhao, X., Tomizuka, M., Yang, X., Yan, J., Ding, M., 2025 · 2025
Closest in time.
Sam2act: Integrating visual foundation model with a memory architecture for robotic manipulation
Fang, H., Grotz, M., Pumacay, W., Wang, Y.R., Fox, D., Krishna, R., Duan, J., 2025 · 2025
Closest in time.
Fu, H., Zhang, D., Zhao, Z., Cui, J., Liang, D., Zhang, C., Zhang, D., Xie, H., Wang, B., Bai, X., 2025 · 2025
Closest in time.
ire-vla: Improving vision-language-action model with online reinforcement learning
Guo, Y., Zhang, J., Chen, X., Ji, X., Wang, Y.J., Hu, Y., Chen, J., 2025 · 2025
Closest in time.
Robocerebra: A large-scale benchmark for long-horizon robotic manipulation evaluation
Han, S., Qiu, B., Liao, Y., Huang, S., Gao, C., Yan, S., Liu, S., 2025 · 2025
Closest in time.
Tla: Tactile-language-action model for contact-rich manipulation
Hao, P., Zhang, C., Li, D., Cao, X., Hao, X., Cui, S., Wang, S., 2025 · 2025
Closest in time.
Plaicraft: Large-scale time-aligned vision-speech-action dataset for embodied ai
He, Y., Weilbach, C.D., Wojciechowska, M.E., Zhang, Y., Wood, F., 2025 · 2025
Closest in time.
Nora: A small open-sourced generalist vision language action model for embodied tasks
Hung, C., Sun, Q., Hong, P., Zadeh, A., Li, C., Tan, U., Majumder, N., Poria, S., et al., 2025 · 2025
Closest in time.
Robobrain: A unified brain model for robotic manipulation from abstract to concrete, in: Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 1724–1734
Ji, Y., Tan, H., Shi, J., Hao, X., Zhang, Y., Zhang, H., Wang, P., Zhao, M., Mu, Y., An, P., et al., 2025 · 2025
Closest in time.
Jiang, S., Li, H., Ren, R., Zhou, Y., Wang, Z., He, B., 2025 · 2025
Closest in time.
Beyond sight: Finetuning generalist robot policies with heterogeneous sensors via language grounding
Jones, J., Mees, O., Sferrazza, C., Stachowicz, K., Abbeel, P., Levine, S., 2025 · 2025
Closest in time.
Clip-rt: Learning language-conditioned robotic policies from natural language supervision, in: Proceedings of Robotics: Science and Systems (RSS)
Kang, G.C., Kim, J., Shim, K., Lee, J.K., Zhang, B.T., 2025 · 2025
Closest in time.
Khan, M., Asfaw, S., Iarchuk, D., Cabrera, M., Moreno, L., Tokmurziyev, I., Tsetserukou, D., 2025 · 2025
Closest in time.
Fine-tuning vision-language-action models: Optimizing speed and success
Kim, M., Finn, C., Liang, P., 2025 · 2025
Closest in time.
Onetwovla: A unified vision-language-action model with adaptive reasoning
Lin, F., Nai, R., Hu, Y., You, J., Zhao, J., Gao, Y., 2025 · 2025
Closest in time.
Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model
Liu, J., Chen, H., An, P., Liu, Z., Zhang, R., Gu, C., Li, X., Guo, Z., Chen, S., Liu, M., et al., 2025 · 2025
Closest in time.
Lykov, A., Serpiva, V., Khan, M.H., Sautenkov, O., Myshlyaev, A., Tadevosyan, G., Yaqoot, Y., Tsetserukou, D., 2025 · 2025
Closest in time.
Myers, V., Zheng, B.C., Dragan, A., Fang, K., Levine, S., 2025 · 2025
Closest in time.
Pre-training auto-regressive robotic models with 4d representations
Niu, D., Sharma, Y., Xue, H., Biamby, G., Zhang, J., Ji, Z., Darrell, T., Herzig, R., 2025 · 2025
Closest in time.
NVIDIA Isaac Sim
NVIDIA Corporation, · 2025
Closest in time.
Gemini robotics on-device brings ai to local robotic devices
Parada, C., Team, G.R., 2025 · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
Pertsch, K., Stachowicz, K., Ichter, B., Driess, D., Nair, S., Vuong, Q., Mees, O., Finn, C., Levine, S., 2025 · 2025
Closest in time.
Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation
Qi, Z., Zhang, W., Ding, Y., Dong, R., Yu, X., Li, J., Xu, L., Li, B., He, X., Fan, G., et al., 2025 · 2025
Closest in time.
Spatialvla: Exploring spatial representations for visual-language-action model
Qu, D., Song, H., Chen, Q., Yao, Y., Ye, X., Ding, Y., Wang, Z., Gu, J., Zhao, B., Wang, D., et al., 2025 · 2025
Closest in time.
Uav-vla: Vision-language-action system for large scale aerial mission generation
Sautenkov, O., Yaqoot, Y., Lykov, A., Mustafa, M., Tadevosyan, G., Akhmetkazy, A., Cabrera, M., Martynov, M., Karaf, S., Tsetserukou, D., 2025 · 2025
Closest in time.
Hi robot: Open-ended instruction following with hierarchical vision-language-action models
Shi, L.X., Ichter, B., Equi, M., Ke, L., Pertsch, K., Vuong, Q., Tanner, J., Walling, A., Wang, H., Fusai, N., et al., 2025 · 2025
Closest in time.
Smolvla: A vision-language-action model for affordable and efficient robotics
Shukor, M., Aubakirova, D., Capuano, F., Kooijmans, P., Palma, S., Zouitine, A., et al., 2025 · 2025
Closest in time.
Reassemble: A multimodal dataset for contact-rich robotic assembly and disassembly
Sliwowski, D., Jadav, S., Stanovcic, S., Orbik, J., Heidersberger, J., Lee, D., 2025 · 2025
Closest in time.
Geomanip: Geometric constraints as general interfaces for robot manipulation
Tang, W., Pan, J.H., Liu, Y.H., Tomizuka, M., Li, L.E., Fu, C.W., Ding, M., 2025 · 2025
Closest in time.
Robobert: An end-to-end multimodal robotic manipulation model
Wang, S., Shan, J., Zhang, J., Gao, H., Han, H., Chen, Y., Wei, K., Zhang, C., Wong, K., Zhao, J., et al., 2025 · 2025
Closest in time.
Diffusion-vla: Generalizable and interpretable robot foundation model via self-generated reasoning
Wen, J., Zhu, M., Zhu, Y., Tang, Z., Li, J., Zhou, Z., Li, C., Liu, X., Peng, Y., Shen, C., Feng, F., 2024 · 2025
Closest in time.
Dexvla: Vision-language model with plug-in diffusion expert for general robot control
Wen, J., Zhu, Y., Li, J., Tang, Z., Shen, C., Feng, F., 2025 · 2025
Closest in time.
From foresight to forethought: Vlm-in-the-loop policy steering via latent alignment
Wu, Y., Tian, R., Swamy, G., Bajcsy, A., 2025 · 2025
Closest in time.
Xu, S., Wang, Y., Xia, C., Zhu, D., Huang, T., Xu, C., 2025 · 2025
Closest in time.
Leverb: Humanoid whole-body control with latent vision-language instruction
Xue, H., Huang, X., Niu, D., Liao, Q., Kragerud, T., Gravdahl, J.T., Peng, X.B., Shi, G., Darrell, T., Sreenath, K., Sastry, S., 2025 · 2025
Closest in time.
Magma: A foundation model for multimodal ai agents, in: Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 14203–14214
Yang, J., Tan, R., Wu, Q., Zheng, R., Peng, B., Liang, Y., Gu, Y., Cai, M., Ye, S., Jang, J., et al., 2025 · 2025
Closest in time.
Robotic control via embodied chain-of-thought reasoning
Zawalski, M., Chen, W., Pertsch, K., Mees, O., Finn, C., Levine, S., 2025 · 2025
Closest in time.
Dexgraspvla: A vision-language-action framework towards general dexterous grasping
Zhong, Y., Huang, X., Li, R., Zhang, C., Liang, Y., Yang, Y., Chen, Y., 2025 · 2025
Closest in time.
Chatvla: Unified multimodal understanding and robot control with vision-language-action model
Zhou, Z., Zhu, Y., Zhu, M., Wen, J., Liu, N., Xu, Z., Meng, W., Cheng, R., Peng, Y., Shen, C., et al., 2025 · 2025
Closest in time.