Fetching the paper…
Reading the bibliography…
Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representations that mismatch continuous perception and control.
Libero: Benchmarking knowledge transfer for lifelong robot learning
Liu, B., Zhu, Y., Gao, C., Feng, Y., Liu, Q., Zhu, Y., and Stone, P · 2023
Earlier work this paper cites.
Learning fine-grained bimanual manipulation with low-cost hardware
Zhao, T., Kumar, V., Levine, S., and Finn, C · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Zitkovich, B., Yu, T., Xu, S., Xu, P., Xiao, T., Xia, F., Wu, J., Wohlhart, P., Welker, S., Wahid, A., et al · 2023
Earlier work this paper cites.
p i _ 0 pi\_0 : A vision-language-action flow model for general robot control
Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., et al · 2024
Earlier work this paper cites.
Octo: An open-source generalist robot policy
Ghosh, D., Walke, H. R., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Kreiman, T., Xu, C., Luo, J., et al · 2024
Earlier work this paper cites.
Training large language models to reason in a continuous latent space
Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., and Tian, Y · 2024
Earlier work this paper cites.
Li, Q., Liang, Y., Wang, Z., Luo, L., Chen, X., Liao, M., Wei, F., Deng, Y., Xu, S., Zhang, Y., et al · 2024
Earlier work this paper cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Jiang, Q., Li, C., Yang, J., Su, H., et al · 2024
Earlier work this paper cites.
Gr00t n1: An open foundation model for generalist humanoid robots
Bjorck, J., Castañeda, F., Cherniadev, N., Da, X., Ding, R., Fan, L., Fang, Y., Fox, D., Hu, F., Huang, S., et al · 2025
Earlier work this paper cites.
Sam 3: Segment anything with concepts
Carion, N., Gustafson, L., Hu, Y.-T., Debnath, S., Hu, R., Suris, D., Ryali, C., Alwala, K. V., Khedr, H., Huang, A., et al · 2025
Earlier work this paper cites.
Cui, C., Ding, P., Song, W., Bai, S., Tong, X., Ge, Z., Suo, R., Zhou, W., Liu, Y., Jia, B., et al · 2025
Earlier work this paper cites.
Graspvla: a grasping foundation model pre-trained on billion-scale synthetic action data
Deng, S., Yan, M., Wei, S., Ma, H., Yang, Y., Chen, J., Zhang, Z., Yang, T., Zhang, X., Zhang, W., et al · 2025
Cited alongside, same era.
Long-vla: Unleashing long-horizon capability of vision language action model for robot manipulation
Fan, Y., Bai, S., Tong, X., Ding, P., Zhu, Y., Lu, H., Dai, F., Zhao, W., Liu, Y., Huang, S., et al · 2025
Cited alongside, same era.
Thinkact: Vision-language-action reasoning via reinforced visual latent planning
Huang, C.-P., Wu, Y.-H., Chen, M.-H., Wang, Y.-C. F., and Yang, F.-E · 2025
Cited alongside, same era.
π 0.5 \pi_{0.5} : A vision-language-action model with open-world generalization
Intelligence, P., Black, K., Brown, N., Darpinian, J., Dhabalia, K., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., et al · 2025
Cited alongside, same era.
Molmoact: Action reasoning models that can reason in space
Lee, J., Duan, J., Fang, H., Deng, Y., Li, B., Liu, S., Fang, B., Zhang, J., Wang, Y. R., Lee, S., et al · 2025
Cited alongside, same era.
Codi: Compressing chain-of-thought into continuous space via self-distillation
Shen, Z., Yan, H., Zhang, L., Hu, Z., Du, Y., and He, Y · 2025
Later among the works it cites.
Ricl: Adding in-context adaptability to pre-trained vision-language-action models
Sridhar, K., Dutta, S., Jayaraman, D., and Lee, I · 2025
Later among the works it cites.
Sim-cot: Supervised implicit chain-of-thought
Wei, X., Liu, X., Zang, Y., Dong, X., Cao, Y., Wang, J., Qiu, X., and Lin, D · 2025
Later among the works it cites.
Softcot: Soft chain-of-thought for efficient reasoning with llms
Xu, Y., Guo, X., Zeng, Z., and Miao, C · 2025
Later among the works it cites.
Deepthinkvla: Enhancing reasoning capability of vision-language-action models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Evaluating real-world robot manipulation policies in simulation
Li, X., Hsu, K., Gu, J., Mees, O., Pertsch, K., Walke, H. R., Fu, C., Lunawat, I., Sieh, I., Kirmani, S., et al · 2025
Cited alongside, same era.
F1: A vision-language-action model bridging understanding and generation to actions
Lv, Q., Kong, W., Li, H., Zeng, J., Qiu, Z., Qu, D., Song, H., Chen, Q., Deng, X., and Pang, J · 2025
Cited alongside, same era.
Ma, X., Xing, L., Zhang, H., Li, W., and Lu, S · 2025
Cited alongside, same era.
Fast: Efficient action tokenization for vision-language-action models
Pertsch, K., Stachowicz, K., Ichter, B., Driess, D., Nair, S., Vuong, Q., Mees, O., Finn, C., and Levine, S · 2025
Cited alongside, same era.
Multimodal chain of continuous thought for latent-space reasoning in vision-language models
Pham, T.-H. and Ngo, C · 2025
Cited alongside, same era.
Reasoning to learn from latent thoughts
Ruan, Y., Band, N., Maddison, C. J., and Hashimoto, T · 2025
Cited alongside, same era.
Qwen3-vl technical report, 2025a
Bai, S., Cai, Y., Chen, R., Chen, K., Chen, X., Cheng, Z., Deng, L., Ding, W., Gao, C., Ge, C., et al
Cited in the paper.
Yin, C., Lin, Y., Xu, W., Tam, S., Zeng, X., Liu, Z., and Yin, Z · 2025
Later among the works it cites.
Robotic control via embodied chain-of-thought reasoning
Zawalski, M., Chen, W., Pertsch, K., Mees, O., Finn, C., and Levine, S · 2025
Later among the works it cites.
Reshaping action error distributions for reliable vision-language-action models
Bai, S., Wang, D., Chi, C., Zhou, W., Lyu, J., Zhao, X., Wang, P., Wang, Z., Xing, L., Zhang, S., et al · 2026
Closest in time.
Starvla: A lego-like codebase for vision-language-action model developing
Community, S · 2026
Closest in time.
Fast-thinkact: Efficient vision-language-action reasoning via verbalizable latent planning
Huang, C.-P., Man, Y., Yu, Z., Chen, M.-H., Kautz, J., Wang, Y.-C. F., and Yang, F.-E · 2026
Closest in time.
Action-sketcher: From reasoning to action via visual sketches for long-horizon robotic manipulation
Tan, H., Co, P., Xu, Y., Rong, S., Ji, Y., Chi, C., Chen, X., Zhang, Q., Zhao, Z., Wang, P., et al · 2026
Closest in time.